One number in place of many
Here are ten delivery times, in minutes, for one restaurant on one evening:
The list is complete and accurate, and it is already hard to work with. Ten more restaurants and you cannot hold it in your head. So you compress: replace the list with one number, or two, and accept that you have thrown information away.
Every trap on this page comes from the same place. Somebody computed a summary that discards exactly the information their question needed.
Where is the centre?
Mean
For the ten delivery times: the total is 310 minutes, divided by 10 orders, so the mean is 31 minutes.
Good for
Anything where the total is the real quantity of interest. Budgets, payroll, total demand, fuel consumed, marks that will be added up anyway. The mean is the only common summary that lets you recover the total: mean × count = total.
Also good for repeated measurements of the same thing, where the errors go both ways and cancel out.
Misleading when
A few extreme values sit far from the rest. The 94-minute order pulls the mean up to 31 minutes, which is longer than nine of the ten actual deliveries. Nobody experienced a typical wait of 31 minutes.
Also misleading for skewed quantities like income, property prices, city populations and file sizes, where a long tail on one side drags the mean away from ordinary experience.
Median
Sorted, the middle two delivery times are 24 and 25, so the median is 24.5 minutes. That describes an ordinary order much better than 31 does.
Good for
"What does a typical one look like?" when the data is skewed or has outliers. Salaries, house prices, delivery times, waiting times, wealth. This is why governments report median household income and not mean household income.
Misleading when
The extremes are the point. Median flood damage across all years is close to zero and tells you nothing about what a flood costs. Median insurance claim hides exactly the claims the insurer must plan for.
The median also cannot be added up. You cannot get a total from it, and medians of subgroups do not combine into a median of everything.
Mode
The mode is the only one of the three that works on things that are not numbers. The most common complaint category, the most common bus route, the most common reason for a refund.
Good for
Categories, and anything you have to stock or plan in discrete units. If you are ordering T-shirts for a college fest, the mean shirt size is meaningless and the mode is what you need. Same for shoe sizes, class timings and the most requested dish.
Misleading when
Several values are nearly as frequent as the winner. A mode of 24 minutes sounds decisive when 24 appeared 11 times and 25 appeared 10 times.
The mode also says nothing about spread or about the total, and on continuous data with no repeats it barely exists.
How spread out is it?
Two bus routes, five days of waiting times each, in minutes.
Bus Reliable
8, 9, 10, 11, 12. Mean 10, median 10.
Bus Unpredictable
0, 2, 10, 18, 20. Mean 10, median 10.
Same mean. Same median. If you have to reach an exam on time, these are not the same bus. Every summary of the centre has thrown away the thing you most need to know, so you need a second number.
Range
Bus Reliable has a range of 4 minutes, Bus Unpredictable 20 minutes. Easy to compute, easy to explain, and it depends on exactly two observations out of however many you collected. One freak value changes it completely. The delivery times have a range of 76 minutes because of a single stuck order.
Use it for a quick sanity check on data you have just received, and for bounded things like the highest and lowest temperature in a day. Do not use it as your main measure of spread.
Interquartile range (IQR)
For the ten delivery times, Q1 is about 22 and Q3 is about 27.5, so the IQR is about 5.5 minutes. The middle half of orders arrived within a five-and-a-half minute window, which is a fair description of the evening. The range said 76.
Good for
Skewed data, and any data where you expect a few extreme values. The IQR ignores the top and bottom quarter entirely, so one stuck delivery cannot move it. It pairs naturally with the median.
Misleading when
The tails matter. An IQR of 5.5 minutes gives no hint that one customer waited 94 minutes. If you are the one answering complaint calls, the tail is your whole job.
Variance and standard deviation
The obvious way to measure spread is to ask how far each value sits from the mean. Try it on Bus Reliable and the deviations are −2, −1, 0, +1, +2, which add up to zero. They always add up to zero, for any dataset, because that is what the mean does. So squaring is used to stop the negatives cancelling the positives.
- Subtract the mean from each value to get its deviation.
- Square each deviation, so that negative and positive misses both count.
- Take the mean of those squares. This is the variance.
- Take the square root, to get back to the original units. This is the standard deviation.
Bus Reliable
Squared deviations: 4, 1, 0, 1, 4. Total 10.
Variance = 10 ÷ 5 = 2 min²
Standard deviation = √2 = 1.4 minutes
Bus Unpredictable
Squared deviations: 100, 64, 0, 64, 100. Total 328.
Variance = 328 ÷ 5 = 65.6 min²
Standard deviation = √65.6 = 8.1 minutes
Read the standard deviation as roughly how far a typical observation sits from the mean. Waits on Bus Reliable are usually within a minute or two of 10. Waits on Bus Unpredictable are usually eight minutes off, in either direction.
Variance is in squared units (minutes squared), which nobody can interpret. It is a step on the way, and it is the form that behaves well in the mathematics, so you will meet it constantly. The standard deviation is the one you report.
Good for
Comparing consistency between two things measured in the same units. Two machines filling bottles to a mean of 500 ml, one with a standard deviation of 1 ml and one with 4 ml: the first is the more consistent machine, and you can say that without knowing anything else about the distributions.
Misleading when
Outliers are present. Squaring gives distant points enormous weight. The delivery times have a standard deviation of 21 minutes, driven almost entirely by one order. With that order replaced by a normal 32 minutes, the standard deviation drops to 4 minutes.
Also when the data is not roughly symmetric. Familiar rules like "about two thirds of values lie within one standard deviation" assume a bell-shaped distribution and quietly fail on skewed data.
Why do some formulas divide by n − 1?
Everything above divides by the number of observations, because we treated those five days as the complete set we wanted to describe. When you have a sample and want to estimate the spread of a much larger population, dividing by n gives an answer that is slightly too small on average, so the standard correction divides by n − 1 instead. Spreadsheets and ChatGPT usually assume the sample version. For five observations the difference is noticeable; for five hundred it is not.
The summary follows the question
Decide what you are asking before you decide what to compute. Almost every wrong summary I see comes from doing it in the other order.
| If the question is… | Report… | And check… |
|---|---|---|
| What is the total, or the per-person share of it? | Mean | Whether outliers or skew make the mean unrepresentative |
| What does a typical case look like? | Median | Whether the extremes need a separate sentence |
| What happens most often? | Mode | Whether the runner-up is nearly as frequent |
| How consistent is it? | Standard deviation, or IQR if skewed | Whether one extreme value is driving the number |
| How bad can it get? | Maximum, or the 95th percentile | How many observations you have in the tail |
| Are these two groups different? | Both centre and spread, for each group | Whether the groups are comparable in the first place |
Report a centre and a spread together, as a habit. A mean on its own is half an answer.
Five ways a correct summary tells a false story
None of these involve arithmetic errors. Every number below is computed correctly.
1. One extreme value moves the mean
Five monthly salaries in a small firm, in thousands of rupees: 28, 31, 34, 36, 40. The mean is 33.8 and the median is 34, so either one describes the firm well.
Now the founder pays herself ₹500k. The salaries are 28, 31, 34, 36, 500. The mean jumps to 125.8 and the median stays at 34. "Average salary at our firm: ₹1.26 lakh a month" is arithmetically true and describes nobody who works there.
2. The groups were mixed together
Two batches take the same two courses. Compare them course by course:
| Batch | Intro course | Advanced course | Overall |
|---|---|---|---|
| Batch X | 80 / 90 = 89% | 5 / 10 = 50% | 85 / 100 = 85% |
| Batch Y | 9 / 10 = 90% | 54 / 90 = 60% | 63 / 100 = 63% |
Batch Y did better in the intro course (90% against 89%) and better in the advanced course (60% against 50%), and worse overall (63% against 85%). Both statements are true. The overall figure is dominated by the fact that Batch Y mostly took the harder course.
This reversal has a name, Simpson's paradox, and it appears whenever a summary is computed across groups of very different sizes or difficulty. The fix is to look for the grouping variable and report within groups.
3. The spread was left out
The two bus routes above. Identical mean, identical median, completely different experience. Any comparison of two groups that reports only centres is incomplete, and you should treat it as incomplete when you read it.
4. Different data, identical summaries
Francis Anscombe built four small datasets in 1973 with nearly the same mean of x, mean of y, standard deviations and correlation. Plotted, one is a straight line, one is a curve, one is a line with a single point dragging it, and one is a vertical stack of points plus a single distant point that creates the entire slope.
The summaries cannot tell them apart. A scatterplot separates them in one second. Plot the data before you summarise it, every time.
5. Counts without denominators
"Mumbai reported four times as many road accidents as Nashik last year." Mumbai also has many times more people and vehicles. A count answers "how many", and the question was almost certainly "how likely", which needs a rate. Whenever you see a raw count being used for comparison, ask what it should have been divided by.
Before you report a summary, or believe one
- Which average is it? If a report says "average" without saying mean or median, find out. On skewed data the two can differ by a factor of four.
- Where is the spread? A centre with no measure of spread is half an answer.
- Have I looked at the data? A histogram or a scatterplot takes ten seconds and catches most of the failures on this page.
- How many observations? A mean of five things and a mean of five thousand deserve very different levels of confidence.
- What is in the tail? Skim the largest and smallest few values. They are either errors you should fix or real cases you should mention.
- Are there groups hiding in here? Split by the obvious ones (course, city, batch, gender) and check whether the story survives.
- Should this be a rate? Counts compared across places or groups of different sizes usually should be.
Notes and sources
- The bus waits, the salaries and the delivery times are invented classroom data, small enough that you can check every number by hand.
- Standard deviations on this page use the population formula (divide by n) so that the arithmetic stays visible. See the note in the spread section.
- F. J. Anscombe, "Graphs in Statistical Analysis", 1973. Four pages, and the four datasets are in it.
- Justin Matejka and George Fitzmaurice, "Same Stats, Different Graphs", 2017. The same trick pushed to a dinosaur.
- Related pages in this course: Principles of data visualization for the plotting side, and Could this just be luck? for deciding whether a difference between two groups is real.