Here's something nobody tells you about a big, quotable number in education research: quoted alone, it reads as proof. It sounds settled, precise, already done.
The honest version of most famous results is smaller than the headline. It arrives with a footnote — naming exactly what was measured, on whom, and how big the study behind it actually was.
This chapter walks two cases that make the same point from different corners of the field. One is a headline tutoring effect whose size turns out to come mostly from a different mechanism than the one advertised. The other is a headline network-science threshold that turns out to describe the wrong kind of structure for the question actually being asked.
In both cases, the footnoted version isn't a downgrade. It's the more useful number — the one you can actually check.
Start with what mastery learning (Bloom, 1968, 1984) actually is: a loop, not a single pass through a unit.
Teach the unit. Test it. Anyone scoring below a high threshold — roughly 80 to 90 percent — gets corrective instruction, then a retest. Only once that threshold is cleared does the class move to the next unit.
That loop changes what the familiar spread of grades across a class actually means. One pass, no correction loop, everyone advancing together regardless of who actually reached the threshold — that is what produces the usual bell curve. It isn't a fixed ceiling built into some learners and not others. It's a description of how the teaching was structured.
Here's the number that made mastery learning famous. Bloom's 1984 studies found one-to-one tutoring beating conventional classroom instruction by roughly two standard deviations — 2.0σ. That's a gap large enough to put the average tutored learner at about the 98th percentile of the untutored group.
Bloom didn't pose this as a settled finding. He posed it as an open search question: individual tutoring for every learner doesn't scale, so what could group instruction do to approach that gap?
The number marked the start of an inquiry. It was never a ceiling already proven to hold in general.
Here's the trap the 2.0σ figure sets, if you stop at the headline: treating it as an established, generally repeatable target — a benchmark any well-designed tutoring or mastery program should be able to approach.
It isn't that. The 2.0σ figure rests on two small, unpublished, three-week studies, using researcher-built tests on novel content — an extreme corner of the space of possible studies.
Far larger bodies of work, done since, have never reproduced an effect that size. The most rigorous published meta-analysis of tutoring finds an average effect closer to g≈0.29 (Nickow, Oreopoulos & Quan, 2024, 89 randomized controlled trials). The largest mastery-learning meta-analysis lands near d≈0.52 on criterion tests (Kulik, Kulik & Bangert-Drowns, 1990). Both are a fraction of the original headline.
Decompose the 2.0σ figure, and it stops being one single tutoring effect at all.
Von Hippel's (2024) analysis attributes roughly 1.1σ of the 2.0σ to the test-and-correct feedback cycle itself — the retest-until-threshold structure mastery learning already supplies, with no tutor present. That leaves tutoring's own independent contribution at closer to 0.9σ.
Counted separately against much larger later evidence, both pieces land far below the original headline. The largest mastery-learning meta-analysis puts the feedback-and-correction effect near d≈0.52 (Kulik, Kulik & Bangert-Drowns, 1990, 103 studies). The most rigorous tutoring meta-analysis puts tutoring's own effect near g≈0.29 (Nickow, Oreopoulos & Quan, 2024, 89 randomized controlled trials).
One variable swings these numbers more than any other: which test measured them.
The identical intervention can read two to six times larger on a test the researchers themselves built than on an independent, standardized test. Cohen, Kulik & Kulik (1982) found tutoring at 0.84σ on researcher-made tests against 0.27σ on standardized tests. Kulik, Kulik & Bangert-Drowns (1990) found mastery learning at 0.52σ on researcher-made criterion tests against roughly 0.08σ on standardized tests.
Test type is the single biggest moderator of how large an effect size looks. An effect size reported with no mention of which kind of test produced it says almost nothing on its own.
The second case comes from a different corner of the field entirely: network science.
A familiar metaphor — percolation — says a structure "connects up" once a random share of its pieces are switched on. Cross a critical fraction, and a single connected mass suddenly spans the whole thing.
The famous specific number attached to that story, p_c ≈ 0.5927, belongs to one exact geometry: site percolation on a two-dimensional square lattice. Leave that exact geometry, and the number carries no meaning at all. A different arrangement of connections has, in general, a different threshold — or no clean threshold whatsoever.
The geometry-free version of this result is both more general and sharper.
Erdős & Rényi (1960) showed that in a random network, a single giant connected component appears suddenly once the average number of connections per point — the average degree, ⟨k⟩ — crosses one.
Below that crossing, a random structure almost surely fragments into many small, disconnected pieces. Above it, one piece suddenly dominates. This crossing, not any grid-specific number, is the real, general reason a thinly-covered space of interconnected things tends to fall apart into islands rather than staying whole.
Now take the opposite case: one fixed, already fully known structure — not a random one whose connections are still being decided.
For that case, the right question isn't a probability threshold at all. It's a direct count of how many separate connected pieces the structure actually has. One piece means the whole thing is coherent. More than one piece means it has split into islands. A piece of size one means something is cut off entirely, connected to nothing.
That count is a standard, deterministic computation — breadth-first or depth-first search over the structure, not a statistical estimate. It applies exactly because the structure in question is already fixed and known, the opposite condition from the one a percolation threshold is built to handle.
Here's the misconception this sets up: assuming a percolation threshold, or its famous specific number, is the right tool to check whether a small, fixed, already fully known structure holds together.
It isn't. A percolation threshold is built for the opposite situation — a structure whose connections are still being randomly decided, where the question is a probability.
When the structure is already fixed and fully known, counting its connected pieces directly is both simpler and the actually correct check. It's a plain, deterministic computation — not a downgrade from a fancier-sounding, probability-flavored number that doesn't even apply to the case at hand.
Pull back, and both cases make the same point from different corners of the field.
A field's most quoted numbers are often an artifact of what was measured. Bloom's 2.0σ, decomposed and re-measured against far larger evidence, both shrinks and splits — a feedback-cycle share and a tutoring share, each near a quarter of the original headline.
Or the number is a metaphor borrowed from the wrong setting. A percolation threshold built for one specific random geometry stands in for a question about one small, fixed, already-known structure — a question a direct connected-components count answers more simply.
In both cases, the sturdier version isn't a compromise. Naming the test type behind an effect size, and counting connected pieces directly instead of reaching for a borrowed threshold, are both more accurate than the headline — and easier to actually check.