[Prompt]
custom: Our episode about life expectancy in times when humans did not expect to live into their 80s was ver | Hosts: corn, herman

[Response]
Corn: Daniel's follow-up to the life expectancy episode. He wants the generalization, not another data point. His question is, if you had to rank the types of statistics most likely to be misread by a general audience, what makes the list. Not individual stats, but the categories. Averages, rates, correlations. He wants the taxonomy.

Herman: And the life expectancy example is the perfect seed for it, because the failure there wasn't that the number was wrong. The number was accurate. Thirty-one years at birth in medieval England. The failure was in what people did with it. They read a population average as an individual prediction, and that one move generates about half of this list.

Corn: The lie lives in the step from number to conclusion.

Herman: That's the whole episode in nine words. And I'd put the mean at the top of the list, because it's the most common and the most quietly deceptive. The mean is a choice, not a measurement. You take nine people in a bar earning forty thousand a year, one person walks in earning four million, the average income in that room jumps to four hundred thirty-six thousand dollars. That figure describes nobody in the room. It's arithmetically true and completely useless.

Corn: The average American family has two point five children. The statistic is fine. The family doesn't exist.

Herman: And the mean gets worse with skewed distributions. Income, wealth, house prices, anything with a heavy tail. The median household income in the US in twenty fourteen was about thirty-three thousand. The mean was about fifty thousand. Same data, two summaries, and they tell completely different stories about whether the typical American is doing fine or struggling.

Corn: Which is why the choice of which average to report is a political act. If you want to make the case that things are getting better, you reach for the mean. If you want to make the case that things are stagnant, you reach for the median. Both are defensible. One of them is honest for skewed data.

Herman: The median is almost always the more honest summary for anything with a long tail. It's the middle value. It doesn't care if the richest person in the sample bought a yacht. But here's the thing that makes this a category of misinterpretation rather than just a preference. Most people don't know which one they're looking at. A headline says average income rose four percent. Is that the mean or the median? The headline doesn't say, and the reader doesn't ask.

Corn: And average itself is a fuzzy word. Statistically it covers mean, median, and mode. Three different numbers, all correctly called the average. So even when a journalist is being careful, the word itself is doing damage.

Herman: Which brings in the mode. The mode is the most common value. In the life expectancy data from Cambridge, the most common age for adult death in England between sixteen hundred and eighteen hundred was around seventy. That's the mode. The mean at birth was thirty-five to forty. Both are averages. One sounds like everyone died young, the other sounds like everyone lived to old age. Same population.

Corn: So number one on the list is the mean, specifically the mean presented without a measure of spread. Number two is the confusion between mean and median. And number three is the word average itself, which lets people smuggle in whichever of the three they prefer.

Herman: I'd put percentage change versus percentage points next. This one shows up constantly in political coverage. A candidate's approval is at fifty percent. It drops twenty percent. Does that mean it's now at forty? Or at thirty? Twenty percent of fifty is ten, so a twenty percent drop takes you to forty. But people say dropped twenty percent when they mean dropped twenty percentage points, which takes you to thirty. The language sounds identical and the numbers are ten points apart.

Corn: I've seen headlines do both in the same week about the same poll.

Herman: And it gets worse when you cross a hundred. If you go from ten tomatoes to fifteen, that's a fifty percent increase. But you now have one hundred fifty percent of your previous yield. Both are true. They sound contradictory. People think the math is broken when it's just two different baselines.

Corn: The percentage change versus percentage points confusion is probably the single most exploited ambiguity in political journalism. It lets a pollster say support fell by twenty percent and have half the audience hear one number and half hear the other.

Herman: Then there's relative versus absolute risk. This is the one that shows up in every health scare. Doubles your risk. Of what, exactly? If your baseline risk is one in a million, doubling it gets you to two in a million. Still tiny. But the headline says doubling, and the reader hears catastrophe. The relative change is meaningless without the baseline it's hiding.

Corn: And the pharmaceutical industry has built entire marketing campaigns on this. A drug that reduces your risk of a rare condition by fifty percent sounds like a miracle. Then you find out the risk went from two in ten thousand to one in ten thousand. The absolute reduction is one in ten thousand. The relative reduction is fifty percent. Both true. Completely different emotional payload.

Herman: The false positive paradox is the next one, and it's the most counterintuitive thing on this list. A test with ninety-nine percent sensitivity and ninety-nine percent specificity for a condition that affects one in a thousand people. Sounds like a great test. You test positive. What's the chance you actually have the condition? About nine percent.

Corn: Nine percent. With a ninety-nine percent accurate test.

Herman: Because the condition is rare. Out of a million people, about a thousand have it, and the test catches nine hundred ninety of them. But it also falsely flags about nine thousand nine hundred ninety healthy people. So the pool of positive results is mostly false positives. The test is fine. The intuition that a positive result means you probably have the condition is the problem.

Corn: The base rate is the thing everyone forgets. The rarer the condition, the more a positive test result needs to be tempered. Doctors know this. The public mostly doesn't. And it's why screening programs for rare conditions generate so much anxiety and so many unnecessary follow-up procedures.

Herman: The blue face paint example is the cleanest version of it. Imagine a perfect test for a rare condition that turns your skin blue. The test is flawless. But if a thousand people in the city are wearing blue face paint for a festival, the test is useless. The false positives swamp the true positives, not because the test is bad, but because the base rate is low and the confound is common.

Corn: Survivorship bias is next. The World War Two bomber story is the canonical case. Planes returning from missions had bullet holes everywhere except the engine and the cockpit. The initial plan was to armor the parts with the holes. Abraham Wald pointed out that the holes showed where planes could take damage and still return. The planes hit in the engine and cockpit never came back. So you armor the parts with no holes, because that's where the fatal damage was.

Herman: And it shows up everywhere. University alumni surveys that report average starting salaries. The graduates who are unemployed or underemployed don't answer the survey. So the average is inflated by the survivors. Mutual fund performance tables. The funds that did badly get closed and disappear from the table, so the surviving funds look better than the industry actually performed.

Corn: The music industry is the same. You hear about the bands that made it. You don't hear about the ten thousand bands that played the same clubs and never got signed. The survivors write the history.

Herman: Simpson's Paradox is the one that feels like a magic trick. UC Berkeley in the nineteen seventies. Men appeared to be admitted at higher rates than women overall. But when you broke it down by department, women were admitted at equal or higher rates in most individual departments. The reversal happened because women applied disproportionately to the more competitive departments. The aggregate hid what was actually going on.

Corn: So the headline was discrimination, and the department-level data said the opposite. The aggregation wasn't wrong. It just answered a different question than the one people thought they were asking.

Herman: And Simpson's Paradox is dangerous because it's not a flaw in the data. It's a flaw in the level of aggregation. You can have two groups where the trend goes one way in each group, and the opposite way when you combine them. Both statements are true. The paradox is real. The mistake is thinking that one of them must be false.

Corn: Correlation versus causation is the one everyone thinks they understand. And mostly they don't. The classic spurious correlation is the pirate one. Global temperatures have risen over the past hundred fifty years, and the number of pirates has declined. Both are caused by industrialization. The correlation is real and meaningless.

Herman: Tyler Vigen's whole project is built on this. Nicolas Cage film appearances correlate with swimming pool drownings. Per capita cheese consumption correlates with people dying by entanglement in bedsheets. The correlations are real. The causation is absurd. And yet the human brain is built to see causation in correlation. We can't help it.

Corn: The over-application cuts both ways, though. There's a reflexive contrarian move where any correlation gets dismissed with a smug correlation isn't causation, even when the mechanism is obvious and the evidence is overwhelming. The phrase becomes a thought-terminating cliché.

Herman: That's the tension. The phrase is correct and overused. It's a shield against bad inference, and a weapon against good inference. The skill is knowing which is which.

Corn: Regression to the mean is the one that explains most of the world's superstitions. A baseball player has a career-best season, gets a huge contract, and the next year he's back to normal. Was he slacking? No. He was always going to drift back toward his average. The extreme performance was partly skill and partly luck, and the luck doesn't repeat.

Herman: It's why punishment seems to work and praise seems to backfire. A student does terribly on a test, gets scolded, and does better next time. The scolding didn't cause the improvement. The terrible performance was an outlier, and the next performance was always going to be closer to average. Same with the student who does brilliantly, gets praised, and comes back down. The praise didn't cause the decline. Regression did.

Corn: And the gambler's fallacy is regression to the mean misunderstood. Monte Carlo, nineteen thirteen. Roulette lands on black twenty-six times in a row. Gamblers lose millions betting on red, because it's due. The wheel doesn't remember. Each spin is independent. The streak doesn't make red more likely. But the intuition that things have to even out is almost impossible to shake.

Herman: The law of small numbers is the little sibling. The US counties with the lowest cancer rates and the highest cancer rates are both small and rural. Small samples swing to extremes. If you flip a coin ten times, getting eight heads isn't remarkable. If you flip it a thousand times and get eight hundred heads, something is wrong with the coin. But people read the ten-flip result as evidence of a loaded coin.

Corn: So the list so far. The mean without spread. Mean versus median. The fuzzy word average. Percentage change versus percentage points. Relative versus absolute risk. The false positive paradox. Survivorship bias. Simpson's Paradox. Correlation versus causation. Regression to the mean and the gambler's fallacy. That's ten.

Herman: And there are honorable mentions. Anscombe's Quartet. Four datasets with identical means, identical variances, identical correlations, and completely different shapes when you graph them. The summary statistics are identical and the data is nothing alike. It's the single best argument for always looking at the scatter plot.

Corn: The inspection paradox. Why your bus is always late and your friends have more friends than you do. You're more likely to arrive at the bus stop during a long gap between buses than a short one, so the wait you experience is longer than the average wait. And your friends having more friends than you is just the same math. Popular people show up in more friend groups, so they're overrepresented in your sample.

Herman: Goodhart's Law. When a measure becomes a target, it stops being a good measure. The McNamara Fallacy. Body counts in Vietnam as a proxy for progress. The measure was accurate and the strategy was wrong. Once you optimize for the number, the number stops meaning what it used to mean.

Corn: The prosecutor's fallacy. The one in a million DNA match doesn't mean there's a one in a million chance the defendant is innocent. It means one in a million people would match by chance. In a city of ten million, that's ten people. The chance the defendant is guilty given the match is one in ten, not one in a million. Sally Clark went to prison on exactly this error.

Herman: That case still bothers me. A pediatrician testifying about sudden infant death used exactly that reasoning. The chance of two cot deaths in one family was calculated as one in seventy-three million. But that's the chance of it happening to any random family. The chance that it happened to this family, given that it happened, is a completely different question. She was convicted on a probability that didn't mean what the jury thought it meant.

Corn: The birthday paradox is the fun one. In a room of twenty-three people, there's a better than even chance that two share a birthday. People think it should be much higher, like a hundred eighty-three. But twenty-three people make two hundred fifty-three pairs, and each pair is a chance at a match. The pairs are what count, not the people.

Herman: So if I had to construct the definitive top ten, the list that most reliably trips people up, I'd order it by frequency of appearance in news coverage times the size of the gap between intuition and reality.

Corn: The mean tops the list because it's everywhere and the failure is invisible. Then mean versus median, because the choice is political and the reader doesn't know it's been made. Then the fuzzy average. Then percentage change versus percentage points, because it's the most exploited ambiguity in political reporting. Then relative versus absolute risk, because health scares run on it.

Herman: The false positive paradox, because it's the most counterintuitive. Survivorship bias, because it's the one that shapes entire industries without anyone noticing. Simpson's Paradox, because it feels like a magic trick and it's real. Correlation versus causation, because everyone knows the phrase and almost nobody applies it correctly. And regression to the mean, because it explains most of the world's superstitions and most of the world's bad management decisions.

Corn: The thing that unites the whole list is that none of these are errors in the data. The numbers are accurate. The failure is in the inference. The step from number to conclusion.

Herman: And that's the uncomfortable part. We teach people to check whether the number is right. We don't teach them to check whether the interpretation is right. A correct number with a wrong interpretation is more dangerous than a wrong number, because the correct number comes with a halo.

Corn: The life expectancy case is the perfect example. The number thirty-one was correct. The interpretation, everyone died at thirty-one, was wrong. And the wrong interpretation survived for centuries because the number was correct and nobody thought to question the step after it.

Herman: The Cambridge data makes the point brutally. In eighteen forty-one, one hundred thirty-eight out of every thousand babies died before age one. Over a quarter were dead by age five. But the children who reached age five had a fifty-fifty chance of reaching sixty. And nearly ten percent of the original thousand reached eighty. Life expectancy that year was forty-two. The mean was accurate. The picture it painted was wrong.

Corn: The most common age for adult death was around seventy. The mode. So the average person didn't die at forty-two. The average person who survived childhood died at seventy. The forty-two was an artifact of infant mortality dragging the mean down.

Herman: And here's the knock-on effect. The myth that everyone died young persists partly because it makes the present feel superior. It's a self-flattering story. Look how far we've come, people used to die at thirty. The reality is more complicated and less comforting. People have always lived to old age when they survived the early hazards. The progress is real, but it's concentrated in infant mortality and maternal death, not in extending the adult lifespan by forty years.

Corn: The mean hid that. The mean at birth collapsed a bimodal distribution into a single number and erased the fact that there were two very different populations in the data. Babies who died in year one, and adults who lived to seventy. The average of those two groups described neither.

Herman: And that's the deeper point about the mean. It's not just that it's sensitive to outliers. It's that with multi-modal data, it can describe a population that doesn't exist. The average of a bimodal distribution is a point between the two modes where there are no people at all.

Corn: The two point five children problem again.

Herman: The average family has two point five children. No family has two point five children. The statistic describes a family that doesn't exist. And if you design policy around that family, you design policy for nobody.

Corn: So the list isn't just a list of statistical errors. It's a list of ways that accurate numbers can be made to lie. And the most common mechanism is the mean, because it's the default. It's what people reach for when they want to summarize data, and it's the one that most reliably misleads when the data is skewed or multi-modal.

Herman: The median is the correction, but the median has its own failure mode. It hides the extremes entirely. If the median income is stagnant but the top one percent is pulling away, the median tells you the typical person is fine and misses the distributional shift. So the median isn't a fix, it's a different tool. The honest summary is usually both, plus a measure of spread.

Corn: The spread is the thing that never makes the headline. Mean income rose four percent. Standard deviation not mentioned. The headline gives you the central tendency and hides the variance. But the variance is where the story is.

Herman: And variance is where the risk lives. Two investments with the same average return. One is a smooth five percent a year. The other is negative twenty percent one year, positive thirty the next. Same mean, wildly different experiences. The mean hides the volatility. The investor who only reads the mean doesn't know what they're buying.

Corn: The false positive paradox is the one I think deserves more attention than it gets, because it's about to get much worse. As we screen for more things, with more sensitive tests, at lower base rates, the false positive problem compounds. The rarer the condition you're screening for, the more likely a positive result is wrong.

Herman: And the screening itself can cause harm. Anxiety, invasive follow-up procedures, overdiagnosis. The test was accurate. The inference was wrong. The patient paid the price.

Corn: The Sally Clark case is the grim endpoint. A woman convicted of murdering her two children because a pediatrician testified that the chance of two cot deaths was one in seventy-three million. The number was arithmetically defensible and completely wrong as an inference. She spent years in prison. The Court of Appeal eventually overturned the conviction, but the damage was done. And she never recovered. She died of alcohol poisoning a few years after release.

Herman: That's the prosecutor's fallacy in its purest form. The probability of the evidence given innocence is not the probability of innocence given the evidence. And the jury couldn't tell the difference, because nobody taught them the difference.

Corn: The Royal Statistical Society actually issued a statement after that case, saying the calculation was wrong. A whole learned society felt compelled to weigh in on a single criminal case because the statistical error was so egregious.

Herman: And it wasn't just the one error. The pediatrician also treated the two deaths as independent events. But cot deaths in the same family are correlated. Genetic factors, environmental factors. The independence assumption was wrong, and the one in seventy-three million figure collapsed.

Corn: The prosecutor's fallacy is on the list, but it's a special case of the broader base rate problem. The base rate is the thing everyone forgets. The prior probability. The context that the number sits in.

Herman: It's the hardest one to teach, because it's counterintuitive. The test is ninety-nine percent accurate. You tested positive. The intuitive answer is you probably have the condition. The correct answer is you probably don't, if the condition is rare enough. The intuition is wrong and the math is right, and that's a hard combination to sell.

Corn: Simpson's Paradox is the one that most undermines trust in statistics. Because it shows that both the aggregate and the disaggregated view can be true, and they can point in opposite directions. When people see that, they conclude that statistics can say anything, and they stop trusting any of it.

Herman: The Berkeley case is the cleanest example. The aggregate showed a bias against women. The department-level data showed no bias, or a bias in the other direction. Both were true. The question was which level of analysis was appropriate for the question being asked. And that's a judgment call, not a math error.

Corn: The danger is that Simpson's Paradox becomes a tool for motivated reasoning. If you don't like the aggregate result, you disaggregate until you find a level where the result flips, and then you declare victory. The paradox is real, but it can be exploited.

Herman: The flip side. Sometimes the aggregate is the right level, and disaggregating is the distortion. If you're looking at whether a university discriminates in admissions overall, the department-level data might hide a real effect in how applicants are steered toward different departments. The choice of level is a substantive question, not a statistical one.

Corn: Survivorship bias is the one that shapes the most decisions without anyone noticing. Every business book about successful companies is a study in survivorship bias. You read about the companies that made it. You don't read about the companies that made the same decisions and failed. So you conclude that the decisions caused the success, when the decisions might have been irrelevant or even harmful, and the survivors survived for other reasons.

Herman: The mutual fund industry is the cleanest example. Funds that underperform get closed or merged. The surviving funds look like they beat the market, because the losers disappeared from the dataset. The published track record is systematically inflated by the missing failures.

Corn: The bomber example is so elegant because the correct answer is the opposite of the intuitive one. The holes showed where the planes could survive damage. The absence of holes showed where the damage was fatal. Armor the absence. It's a complete inversion of the naive read.

Herman: Wald's insight was that the data you don't see is sometimes more informative than the data you do see. The planes that didn't return were the data. And they were missing from the sample by definition.

Corn: That's the unifying theme of the whole list. The most important data is often the data you can't see. The babies who died before the census. The graduates who didn't answer the survey. The funds that got closed. The planes that didn't return. The false positives you never hear about because they got a clean bill of health. The absence is the story.

Herman: The mean is the tool that most reliably hides the absence. It summarizes what's there and ignores what's not there. The life expectancy at birth averaged in the babies who died and the adults who lived, and produced a number that described neither. The absence of the distinction between those two populations was the error.

Corn: The top ten list is really a list of ways that summary statistics hide the structure of the data. The mean hides skew and multimodality. The median hides extremes. The percentage change hides the baseline. The relative risk hides the absolute risk. The false positive rate hides the base rate. Survivorship bias hides the missing cases. Simpson's Paradox hides the subgroup structure. Correlation hides the mechanism. Regression to the mean hides the luck.

Herman: The common thread is that all of these are fixable with more information. The mean plus the spread. The relative risk plus the baseline. The correlation plus the mechanism. The aggregate plus the disaggregation. The number plus the context.

Corn: The problem is that the context is the first thing to go in a headline. The headline has room for the number and maybe one qualifier. The context gets cut. And the reader is left with a number that is technically accurate and practically misleading.

Herman: The News Literacy Project has a whole curriculum on this. They use the phrase numbers don't lie, but even accurate numbers can paint a misleading picture. The lie lives in the step from number to conclusion.

Corn: The actionable version of the list. When you see a headline with an average, ask which average. When you see a percentage change, ask what the baseline is. When you see a risk, ask whether it's relative or absolute. When you see a test result, ask what the base rate is. When you see a success story, ask who didn't make it into the sample. When you see a correlation, ask what the mechanism is. When you see an extreme performance, expect regression.

Herman: That's the checklist. It's not hard. It's just not taught. And the cost of not knowing it is enormous. Bad medical decisions, bad investment decisions, bad policy decisions, wrongful convictions. The list of harms from statistical misinterpretation is not abstract.

Corn: The Sally Clark case is the one that sticks with me. A woman lost her children, then lost her freedom, then lost her life, because a pediatrician and a jury didn't understand the difference between the probability of the evidence given innocence and the probability of innocence given the evidence. That's not a rounding error. That's a tragedy.

Herman: The tragedy is that the error was preventable. The math isn't hard. The concepts aren't beyond a general audience. They're just not taught. We teach algebra and geometry and trigonometry, and we don't teach the difference between relative and absolute risk. We send people into the world with the ability to solve quadratic equations and no ability to read a health headline.

Corn: The quadratic equation has killed far fewer people than the false positive paradox.

Herman: The list, in order. The mean without spread. Mean versus median confusion. The fuzzy word average. Percentage change versus percentage points. Relative versus absolute risk. The false positive paradox. Survivorship bias. Simpson's Paradox. Correlation versus causation. Regression to the mean and the gambler's fallacy.

Corn: The honorable mentions. Anscombe's Quartet. The inspection paradox. Goodhart's Law. The prosecutor's fallacy. The birthday paradox.

Herman: The birthday paradox is the one I'd use to teach the whole list. It's fun, it's counterintuitive, and it's true. Twenty-three people, fifty percent chance of a shared birthday. The intuition says it should be much higher. The math says the pairs are what count. And once you understand why, you understand something about how probability actually works.

Corn: The inspection paradox is the one that explains the most everyday frustration. Your bus is always late because you're more likely to arrive during a long gap. Your friends have more friends than you because popular people are overrepresented in friend groups. The paradox is real and the explanation is structural, not personal.

Herman: Goodhart's Law is the one that explains the most institutional dysfunction. Once a measure becomes a target, it stops measuring what it was meant to measure. School rankings. Hospital wait times. Police arrest quotas. The measure is accurate and the behavior it induces is destructive.

Corn: The McNamara Fallacy is the canonical case. Robert McNamara ran the Vietnam War on body counts. The measure was accurate. The strategy was wrong. The war was lost on accurate data.

Hilbert: We ran a body count on a factory I worked at in seventy-nine. Counted the number of parts that came off the line each shift. Foreman got a bonus tied to the count. Within a month the parts were coming off the line faster and the reject bin was overflowing. The count was accurate. The parts were garbage. They scrapped the bonus and the count went back down. Nobody ever mentioned the reject bin in the report. The number didn't lie. It just didn't count what mattered.

Corn: The reject bin is the survivorship bias in miniature. The parts that failed didn't make it into the count.

Herman: The foreman optimized for the number, which is Goodhart's Law. The measure became the target, and the target stopped measuring quality.

Hilbert: We had a phrase for it on the floor. Counting the wrong thing faster. That's what the bonus did. We counted the wrong thing faster.

Corn: Counting the wrong thing faster. That's the whole list in five words.

Hilbert: I saw the same thing in a warehouse job years later. They timed how fast you picked orders. The fast pickers were the ones who grabbed the wrong items and let the customer sort it out. The slow pickers were the ones who checked the labels. The metric rewarded the fast pickers. The customers paid for it.

Herman: The metric was accurate and the behavior it induced was destructive. Same pattern.

Hilbert: The warehouse manager couldn't see it because the number looked good. Orders were going out faster. The returns data was in a different system. Nobody put the two together. The number looked good and the business was bleeding.

Corn: The absence is the story again. The returns data was the missing cases. The planes that didn't return.

Hilbert: I told the manager once. He said the numbers don't lie. I said the numbers don't tell you about the returns. He said that's a different department. I let it go.

Herman: The numbers don't lie is the problem. The numbers are fine. The inference is the problem.

Hilbert: I still think about that warehouse sometimes. The fast pickers got promoted. The careful ones got fired. The numbers said the fast ones were better. The numbers were wrong about what better meant.

Corn: The careful ones were the data you couldn't see. Their accuracy didn't show up in the speed metric. The fast pickers' errors didn't show up either. Both were invisible to the number.

Hilbert: That's the thing about numbers. They only see what you point them at.

Herman: Pointing is a choice. The metric is a choice. The level of aggregation is a choice. The baseline is a choice. The number is accurate and the choice is invisible.

Corn: The one thing I'd take from this episode is that the most dangerous statistical error is not a wrong number. It's a correct number with a missing context. The mean without the spread. The relative risk without the baseline. The positive test without the base rate. The success story without the survivors. The number is fine. The inference is the problem.

Herman: The fix is a habit, not a formula. When you see a summary statistic, ask what it's summarizing and what it's leaving out. The absence is usually the story.

Corn: Thanks to our producer Hilbert Flumingtop.

Herman: This has been My Weird Prompts, the human-AI collaboration podcast.

Corn: If you enjoyed this, leave us a review wherever you listen. We'll be back soon.