The one idea
A number can be completely accurate and still tell you the wrong thing. The distortion almost never lives in the arithmetic. It lives in what got counted, who got counted, and which two moments were compared.
What is this a number of, and who is missing from it?
Before believing a number, ask what it counted and who got left out.
Check base, denominator, time window and selection before repeating any figure in a document with your name on it.
Most misleading statistics are sound arithmetic over an unrepresentative sample, an unstated denominator, or a selected interval. The error is in the framing, not the calculation.
The eight
1. A percentage with no base
"Sign-ups rose 200%." From two to six. Percentages compress size out of existence, which is why small companies quote growth rates and large ones quote absolute numbers.
Ask: 200% of what starting number?
2. An average hiding the distribution
Ten employees. Nine earn ₹5 lakh, one earns ₹55 lakh. The average is ₹10 lakh — a figure describing nobody in the room. Averages are safe when things cluster and deceptive when they do not, which covers salaries, response times, revenue per customer, and most things with a long tail.
Ask: what is the median, and what does the spread look like? If someone reports an average response time of four hours, the customers who waited three days are still in there somewhere, invisible.
3. A cherry-picked timeframe
Any series that moves contains a window where it rose and one where it fell. Choosing where the line starts produces a story without saying anything false.
Ask: why does the chart start there? If a graph begins at an odd month, that month was probably a low point.
4. Survivorship bias
You read that most successful founders dropped out of college — but you are only counting successful founders. The dropouts whose companies died are not in the dataset, because failure does not get interviewed.
At work: "everyone who took the training got promoted" (measured among people still employed), or "our customers love the product" (measured among people who did not leave).
Ask: who left the dataset before it was counted?
5. A small sample
Five people is a conversation, not a finding. The trap is that small samples produce more extreme results — so the most dramatic number in a report is often the one from the smallest group.
Ask: how many? And if a report breaks results down by segment, how many are in each segment?
6. Correlation dressed as causation
Two things move together. That is compatible with A causing B, B causing A, something else causing both, or coincidence. Business writing picks the first and states it as established.
The third is commonest. Teams using the new project tool ship faster — and are also the teams organised enough to adopt a new tool.
Ask: what third thing could produce both?
7. A truncated axis
A bar chart whose axis starts at 90 rather than 0 turns a one-point change into a cliff, and dashboards do this by default. Not always dishonest — sometimes the interesting variation is in a narrow band — but it must be read for.
Ask: where does the axis start, and what does the change look like against the full scale?
8. Relative risk without absolute risk
"This doubles your risk." From what to what? From 1 in 100,000 to 2 in 100,000, "doubles" is correct and irrelevant. Relative figures alarm; absolute figures let you decide. Same shape at work: "reduces errors by 50%" means one thing if you had 400 errors a month and another if you had four.
Ask: what were the two actual numbers?
The mental model
| The claim | The question | Why it works |
|---|---|---|
| A percentage | Of what base? | Percentages hide size |
| An average | What is the median? | Averages hide spread |
| A trend | Why does it start there? | Windows are chosen |
| A success rate | Who dropped out? | Failures leave the data |
| A comparison | How many in each group? | Small groups swing wildly |
| "X leads to Y" | What else causes both? | Third variables are everywhere |
| A chart | Where is zero? | Axes are choices |
| "Doubles the risk" | From what to what? | Relative hides absolute |
A vendor pitch: “Companies using our platform see 3x faster onboarding.”
Base: 3x what, measured how? Selection: companies buying an onboarding platform had an onboarding problem and decided to fix it, which involves rewriting the process. Survivorship: this is measured on customers who stayed — those who abandoned it in month two are not in the number.
You do not need to conclude the product is bad. You need one question in the call: “Is that measured across all customers who bought it, or across current customers?” The answer, and how comfortable they look giving it, tells you most of what you need.
Your own dashboard says average ticket resolution is 6 hours, comfortably inside the 8-hour target. Nobody is complaining internally, but customers are.
Check the median: if it is 40 minutes against a 6-hour average, a few very slow tickets are doing all the damage, and those are the angry customers. Then check what counts as "resolved" — if an auto-close after 7 days of silence counts, your worst outcomes are recorded as successes.
Both are ordinary reporting decisions made by reasonable people, which is why this is worth checking rather than accusing anyone.
Try this
A colleague writes:
“Since we introduced the new hiring process in March, the quality of hires has clearly improved — 85% of new joiners rated the process positively, and attrition among people hired since March is only 3%, compared to 12% company-wide.”
Find at least three problems before reading on.
Checking a number in four minutes
1 of 4Find the original source, not the article quoting it. Search the exact phrase in quotation marks. Repetition is how a weak number becomes an accepted one — the fifth article citing it does not add evidence.
Your challenge
Level 3 · IndependentTake a chart or figure from a real dashboard, news article, or company deck you have access to. Produce a short note containing:
- The number as stated.
- Its denominator, or a note that you could not find one.
- Who is excluded from the measurement.
- The number rewritten in absolute terms.
- One sentence on whether it still supports the conclusion it was used for.
Success criterion: someone reading your note can decide whether to trust the original without going back to it.
What people usually get wrong
- Trusting a number because it is precise. "37.4%" feels more rigorous than "about a third". Precision is a formatting choice, not evidence.
- Trusting a number because it came from a dashboard. Dashboards are built by people who made definitional choices you cannot see.
- Comparing two numbers produced by different methods. Changing survey tool, vendor, or definition mid-year breaks the comparison, and the break is rarely labelled on the chart.
- Treating "no change" as no finding. A flat line after a big investment is a result, and an important one.
- Becoming the person who distrusts all numbers. This is not scepticism, it is a different way of avoiding thinking. The goal is to know which numbers to lean on.
How someone experienced does it
Experienced people ask for the denominator before the headline figure, and they ask about the measurement definition before either. "Before we look at the number — what counts as an active user?" changes the entire conversation, and often ends it, because frequently nobody knows.
They also watch for the number that did not appear. A report giving you conversion rate but not volume, or growth rate but not base, has made a choice about what you get to see. The missing number is usually the interesting one.
And when presenting numbers themselves, they state the definition and the exclusion on the same slide: "Median, not mean. Excludes the two enterprise accounts, which would move it by 3 points." That is credibility insurance — nobody can later discover something you already told them.
When not to use this
Do not run this audit on every number you encounter. Most numbers at work are routine operational figures where the stakes of being slightly wrong are near zero.
Spend the effort where a number is being used to justify a decision, where it is going into something external, or where it is surprisingly good news. The last one matters most: numbers that flatter the person presenting them get checked least, which is exactly backwards.
Why the smallest group always has the most dramatic result
When a report breaks results down — by region, team, product line — the segment with the most extreme number is very often the one with fewest people in it.
This is arithmetic, not conspiracy. In a small group one unusual individual moves the average a long way; in a large group they are absorbed. "Our Pune office has 40% higher engagement" may mean nine people, two of them enthusiastic.
The consequence is a management failure that repeats everywhere: someone picks the most striking segment out of a breakdown and builds a policy on noise. Whenever a segment leaps out, look at its count before its value.
Prove it
Find one statistic repeated inside your organisation or field — the kind everyone quotes and nobody sources. Track it back to its origin.
Write a paragraph on what it actually measured. Often the trail ends at a source saying something narrower than what everyone repeats, and you will be the only person who knows.
Keep learning this
Paste this into any AI assistant. It turns the assistant into a tutor that tests you instead of just answering you.
Act as an experienced practitioner who is good at teaching. I have just learned checking whether a specific statistic supports the claim it is used for. Assume I am intelligent but relatively new to this — treat me as intermediate level. Work through this in order, and wait for my reply at each step: 1. Ask me 5 questions that test whether I actually understood checking whether a specific statistic supports the claim it is used for. Do not reveal the answers yet. 2. After I answer, tell me which parts I got right, which I got wrong, and which I only half-understand. Explain only what I misunderstood — do not re-teach what I already know. 3. Give me one practical challenge based on something I could genuinely encounter at work or in daily life. Do not solve it for me. 4. Evaluate my solution the way an experienced person would judge it, including what a professional would have done differently. 5. Tell me what to learn next, and why that comes next. 6. Give me trustworthy sources for deeper study — prefer official documentation, primary research or standards bodies over blogs and videos. Rules for you: no buzzwords. No motivational filler. Say "I'm not certain" when you are not certain, and tell me which parts of your answer I should verify myself. Clearly separate facts from your recommendations and your opinions.
Become independent at this
Use this when you want a path from where you are to actually good, with checkpoints you can test yourself against.
I want to become independently capable at reading statistics critically — not permanently dependent on AI, tutorials or step-by-step guides. Design a progression for me with five stages: Beginner, Guided practice, Independent practice, Real-world application, Professional level. For each stage tell me: - what I must know - what I must be able to do without help - the mistakes people make at this stage - one practical challenge - one real project that would prove I reached this stage - one way I can test myself honestly Then tell me the signals that I am ready to move to the next stage, and the signals that I have skipped ahead too early. Keep the theory to the minimum I actually need. Focus on ability I can transfer to situations you and I have not discussed.