Sampling: tasting the soup
How a small random sample can describe millions of people, and why a huge sample picked the wrong way can't.
WHO YOU ASK
Population and sample
The population is everyone you want to know about: all adults in a country, all customers of a shop, all the light bulbs a factory made this week. The sample is the part you actually measure. You almost never measure the whole population, because it costs too much or takes too long. So the real question is always: does this sample look like the population?
A pollster wants to know what all 60 million adults in a country think, and phones 1,000 of them. The 60 million are the population; the 1,000 are the sample.
One spoonful of a stirred pot
You don't drink the whole pot to check the salt; one spoonful tells you. But only if the soup has been stirred. Taste from the top of an unstirred pot and you might get all broth and no salt, however big your spoon is. A sample works the same way: a small one can describe a huge population if it is picked so it mixes everyone in fairly.
Check yourself
A university wants to know how satisfied its 20,000 students are, and emails a survey to 500 of them. What is the population here?
- The 500 students who got the email
- All 20,000 students at the university
- Every student in the country
- The students who answer the email
Show the answer
All 20,000 students at the university
Right. The population is everyone the university wants to know about: all 20,000 of its students. The 500 are the sample, and the ones who answer are a smaller group still.
Step through it

A town of 1,000 people A town of 1,000 people, one dot each, labelled N = 1000. The orange dots, 400 of them or 40 in every 100, prefer tea; the grey ones don't. The empty box on the right is waiting for a sample.

A random 100: the sample says 41% One hundred people scattered all over the town are ringed in lilac and lifted into the box, labelled n = 100. Their answers are stacked orange first: 41% of the sample prefer tea, close to the true 40%.

New random samples, new answers The rings keep jumping to a new random 100, and the box reads 38%, then 43%, then 40%, then 41%. Each random sample lands a little differently, but always near 40%. That wobble is called sampling error.

Ask only one corner: 70% Now the 100 are all taken from one corner, the top-left block of the grid, where tea drinkers happen to crowd together. The box reads 70%, in orange, when the truth is 40%. This is a biased sample, and asking more people from the same corner wouldn't fix it.
Check yourself
A news site's online poll gets 50,000 votes. A polling firm randomly samples 1,000 adults. Which better describes what the country thinks?
- The news site's poll, because 50,000 is far more people
- The random 1,000, because everyone had a fair chance of being picked
- Both equally, since any poll over 1,000 people is reliable
Show the answer
The random 1,000, because everyone had a fair chance of being picked
Right. The 50,000 are the site's readers who chose to vote, a corner of the country, like the top-left of the grid. A random 1,000 is a stirred spoonful.
Random sampling
Picking the sample so that every member of the population has a known, ideally equal, chance of being chosen. That's the statistical version of stirring the pot. The sample then tends to look like the population, and the leftover errors are small chance wobbles, called sampling error, whose size can be predicted. Lesson 5 measures it.
Random samples of 100 from our town gave 38%, 43%, 40% and 41% when the truth was 40%: never exact, always close.
Check yourself
What made the 1936 Literary Digest poll wrong, despite millions of replies?
- Too few people replied
- A biased sample: who was on its lists, and who chose to reply
- Roosevelt's voters lied to the magazine
- The magazine added up the votes wrong
Show the answer
A biased sample: who was on its lists, and who chose to reply
Right. Its lists leaned towards better-off people, and those who replied differed again. Millions of answers from a tilted corner still give a tilted answer.
Self-selection
When people choose for themselves whether to be in the sample. Online polls, phone-in votes and "rate us" surveys hear mostly from the most motivated, delighted or annoyed. Their results describe the people who answered, not the public. This is one kind of sampling bias: any method that favours some people over others.
A restaurant's review page is full of 1-star and 5-star ratings. The many diners who thought it was fine rarely bother to write anything.
Check yourself
If a sample is biased, doubling its size removes the bias.
Show the answer
False
False. More people picked the same way just repeats the same tilt more precisely. Asking 200 people from the tea-heavy corner still gives about 70%, not 40%. Only changing how people are picked fixes bias.
Two more biases worth naming
Non-response
The people who don't answer differ from those who do. If busy parents and night-shift workers never pick up the phone, a survey about free time will hear from the people who have the most of it.
Survivorship
You only see the ones that survived. "Old buildings were built better" ignores all the old buildings that fell down or were torn down; only the sturdy ones are still standing to be admired.
Check yourself
A gym surveys its own members, finds that 90% exercise every week, and announces "90% of people in this city exercise weekly". What's wrong?
- Nothing, 90% is a clear result
- The sample is gym members, not the city's population it claims to describe
- The gym should have asked fewer members
- Weekly exercise can't be measured by a survey
Show the answer
The sample is gym members, not the city's population it claims to describe
Right. Gym members are exactly the people most likely to exercise. The result may be true of members, but the population the claim is about, everyone in the city, was never sampled.
Four questions for any survey
- Who was asked?
Does the group match the population the headline talks about, or is it one corner of it?
- How were they chosen?
At random, or did they volunteer, click a link or happen to be nearby?
- How many answered?
A low response rate means the answers may come from an unusual slice of the people asked.
- Who is missing?
People without internet, people who didn't survive, people too busy to reply. Could they have answered differently?
Lesson recap
- The population is everyone you want to know about; the sample is the part you actually measure.
- A random sample is a stirred spoonful: in a town where 40% prefer tea, random samples of 100 landed near 40% every time.
- Sampling bias, like asking only one corner, gave 70%, and asking more people the same way doesn't fix it.
- The 1936 Literary Digest poll had about 2.4 million replies and still picked the wrong winner; size did not beat bias.
- Watch for self-selection, non-response and survivorship, and ask who was asked, how, how many answered and who is missing.