Pathwise

Statistics and Data Literacy · Lesson 7 of 12 · 12 min

Base rates: the number people forget

Why a positive result on a '90% accurate' test can still mean you are probably fine, worked out on a grid of 1,000 people.

Base rate

NOUN · STATISTICS

How common something is before you look at any evidence about one particular person or case. If a condition affects 1% of people, its base rate is 1 in 100. Every test result has to be read against it.

In a town of 1,000 people where 1% have a condition, the base rate tells you to expect about 10 people with it, before anyone is tested.

TWO WAYS A TEST CAN BE WRONG

Catching the ill, and flagging the healthy

A test's quality has two sides. Sensitivity is how many of the people who really have the condition test positive: a test with 90% sensitivity catches 9 of every 10. A false positive is a healthy person the test wrongly flags. No real test is perfect on both sides, and when you screen lots of healthy people, even a small false-positive rate adds up to a lot of false alarms.

A test that catches 9 in 10 ill people and wrongly flags about 9 in 100 healthy people can fairly be called "about 90% accurate". Keep those two numbers; we'll use them on a grid of 1,000 people.

Check yourself

A condition affects 1% of adults in a country. What is its base rate?

  1. 90%, because that's how accurate the test is
  2. 1 in 100: how common the condition is before anyone is tested
  3. 9%, the chance a positive result is right
  4. It can't be known until everyone is tested
Show the answer

1 in 100: how common the condition is before anyone is tested

Right. The base rate is just how common the condition is in the group, here 1 in 100. It comes before any test and has nothing to do with the test's accuracy.

Step through it

  1. 1,000 people, 10 with the condition

    A grid of 1,000 dots, one for each person in a town. Ten of them, scattered through the grid, are orange: those are the people who have the condition, 1 in 100. The other 990, in grey, are healthy.

  2. The test catches 9 of the 10

    Everyone takes the test. Nine of the ten orange dots get a lilac ring, which marks a positive result: the test correctly catches 9 of the 10 people who are ill, its 90% sensitivity. One orange dot, lower in the grid, has no ring: that ill person is missed.

  3. It also flags about 89 healthy people

    Now 89 grey dots, scattered all over the grid, get lilac rings too. These are false positives: healthy people the test wrongly flags, about 9% of the 990. There are so many healthy people that a small error rate makes a big crowd, and the ringed grey dots far outnumber the ringed orange ones.

  4. Of 98 positives, only 9 are ill

    Everyone who tested positive is gathered into a lilac box on the right: a row of 9 orange dots and 89 grey ones, 98 positive results in all. Only 9 of those 98 people are ill. So in this town a positive result means about a 9% chance of having the condition, not 90%.

Check yourself

In the town of 1,000, the test flagged 9 of the 10 ill people and 89 of the 990 healthy ones. Arash tests positive. About how likely is it that he has the condition?

  1. About 90%
  2. About 50%
  3. About 9%, since 9 of the 98 positives are ill
  4. About 1%, the same as before the test
Show the answer

About 9%, since 9 of the 98 positives are ill

Right. There are 9 + 89 = 98 positive results, and only 9 of them belong to ill people. 9 out of 98 is about 9%, roughly 1 in 11. The positive result did raise his chance, from 1% to about 9%, but nowhere near 90%.

THE MISTAKE HAS A NAME

Base rate neglect

Jumping from "the test is 90% accurate" to "I'm 90% likely to be ill" is called base rate neglect: judging the evidence while ignoring how rare the thing was to start with. The rule of thumb is simple. The rarer the condition, the larger the share of positives that are false, even for a good test, because the healthy crowd is so much bigger than the ill one.

It's not that the test is bad. It did its job: it caught 9 of 10 ill people. The surprise comes entirely from the 990 healthy people it was also run on.

Check yourself

If a test is 90% accurate, a positive result means a 90% chance you have the condition.

Show the answer

False

False. What a positive means depends on the base rate too. For a condition 1 in 100 people have, the same test's positives were right only about 9% of the time, because 89 false alarms from the healthy majority swamped the 9 real cases.

Same test, two different groups

Screening everyone (1% ill)

1,000 people, 10 ill. The test flags 9 of them and about 89 healthy people.

Of 98 positives, 9 are real: about 9%.

People with symptoms (50% ill)

1,000 people, 500 ill. The test flags 450 of them and about 45 of the 500 healthy.

Of 495 positives, 450 are real: about 91%.

Check yourself

The same test is used in a clinic where half the patients have the condition. Compared with screening the whole town, is a positive result there more or less trustworthy?

  1. More trustworthy: when the condition is common, most positives are real
  2. Less trustworthy: more ill people means more errors
  3. Exactly the same, because it's the same test
  4. It depends only on the test's sensitivity
Show the answer

More trustworthy: when the condition is common, most positives are real

Right. The test didn't change; the group did. With half the patients ill, real cases far outnumber false alarms, and about 91% of positives are true.

Check yourself

  1. A bank's fraud alarm catches 95% of fraudulent payments, but only 1 payment in 10,000 is fraud. Most customers whose card is blocked turn out to have done nothing wrong.
  2. The same alarm is used on a small set of accounts already reported as hacked, where about half the payments are fraud. Most blocked payments there really are fraud.

What explains the difference between these two cases?

  1. The alarm works better on hacked accounts
  2. How rare fraud is in each group: when it's rare, most alarms are false
  3. Customers in the first group are more honest
  4. The first bank set its alarm too sensitively
Show the answer

How rare fraud is in each group: when it's rare, most alarms are false

Exactly. It's the same alarm with the same accuracy. What differs is the base rate. Spam filters, airport security and face recognition in crowds all run into the same thing: when what you're looking for is rare, most alarms are false.

When you hear about a test or an alarm

  • Ask first: how common is the thing in the group being tested?
  • Turn percentages into people: imagine 1,000 and count the real and false positives.
  • Rare condition plus mass screening means many positives are false, even for a good test.
  • A positive is a reason for a second look, not a verdict.

Lesson recap

  • The base rate is how common something is before any evidence; here 10 of 1,000 people, 1%.
  • A test that catches 9 of the 10 ill people also flagged about 89 of the 990 healthy ones.
  • Of the 98 positives only 9 were ill, so a positive meant about a 9% chance, not 90%.
  • Ignoring how rare the condition is, base rate neglect, is the classic mistake; counting people on a grid of 1,000 avoids it.
  • Where the condition is common, the same test's positives are mostly real: the group changes the answer.

Keep it, don't just read it

Pathwise brings each idea back just before you'd forget it, with a quick question. Free on Android and on the web, in English and Persian.

Cafe Bazaar Myket Open the web app

All lessons in this course

  1. Mean, median and mode: three kinds of average
  2. Spread: what the average leaves out
  3. The bell curve
  4. Sampling: tasting the soup
  5. Margin of error: the plus-or-minus
  6. Correlation is not causation
  7. Base rates: the number people forget
  8. Percent or percentage points? Relative and absolute risk
  9. Charts that mislead
  10. P-values in plain words
  11. Reading a study
  12. Putting it together: reading a headline