What the big claims are actually worth

FRAME

Here's something nobody tells you about a big, quotable number in education research: quoted alone, it reads as proof. It sounds settled, precise, already done.

The honest version of most famous results is smaller than the headline. It arrives with a footnote — naming exactly what was measured, on whom, and how big the study behind it actually was.

This chapter walks two cases that make the same point from different corners of the field. One is a headline tutoring effect whose size turns out to come mostly from a different mechanism than the one advertised. The other is a headline network-science threshold that turns out to describe the wrong kind of structure for the question actually being asked.

In both cases, the footnoted version isn't a downgrade. It's the more useful number — the one you can actually check.

the headline number vs. the footnoted one
SURFACE
The headline numberquoted alone, it sounds like settled proof — a huge effect, a famous threshold
the footnote is most of the story
BENEATH
The honest, footnoted numbersmaller, or the wrong metaphor entirely, once what was actually measured, and where, is named

↑ Back to top

Mastery learning: the loop, not a single pass

KEY-TERM

Start with what mastery learning (Bloom, 1968, 1984) actually is: a loop, not a single pass through a unit.

Teach the unit. Test it. Anyone scoring below a high threshold — roughly 80 to 90 percent — gets corrective instruction, then a retest. Only once that threshold is cleared does the class move to the next unit.

That loop changes what the familiar spread of grades across a class actually means. One pass, no correction loop, everyone advancing together regardless of who actually reached the threshold — that is what produces the usual bell curve. It isn't a fixed ceiling built into some learners and not others. It's a description of how the teaching was structured.

the loop, not a single pass
Teach the unitTest itCorrective instruction below ~80-90%RetestAdvance

↑ Back to top

The two-sigma story, as posed

CONCEPT

Here's the number that made mastery learning famous. Bloom's 1984 studies found one-to-one tutoring beating conventional classroom instruction by roughly two standard deviations — 2.0σ. That's a gap large enough to put the average tutored learner at about the 98th percentile of the untutored group.

Bloom didn't pose this as a settled finding. He posed it as an open search question: individual tutoring for every learner doesn't scale, so what could group instruction do to approach that gap?

The number marked the start of an inquiry. It was never a ceiling already proven to hold in general.

MISCONCEPTION

Here's the trap the 2.0σ figure sets, if you stop at the headline: treating it as an established, generally repeatable target — a benchmark any well-designed tutoring or mastery program should be able to approach.

It isn't that. The 2.0σ figure rests on two small, unpublished, three-week studies, using researcher-built tests on novel content — an extreme corner of the space of possible studies.

Far larger bodies of work, done since, have never reproduced an effect that size. The most rigorous published meta-analysis of tutoring finds an average effect closer to g≈0.29 (Nickow, Oreopoulos & Quan, 2024, 89 randomized controlled trials). The largest mastery-learning meta-analysis lands near d≈0.52 on criterion tests (Kulik, Kulik & Bangert-Drowns, 1990). Both are a fraction of the original headline.

↑ Back to top

Decomposing the 2.0 sigma

CONCEPT

Decompose the 2.0σ figure, and it stops being one single tutoring effect at all.

Von Hippel's (2024) analysis attributes roughly 1.1σ of the 2.0σ to the test-and-correct feedback cycle itself — the retest-until-threshold structure mastery learning already supplies, with no tutor present. That leaves tutoring's own independent contribution at closer to 0.9σ.

Counted separately against much larger later evidence, both pieces land far below the original headline. The largest mastery-learning meta-analysis puts the feedback-and-correction effect near d≈0.52 (Kulik, Kulik & Bangert-Drowns, 1990, 103 studies). The most rigorous tutoring meta-analysis puts tutoring's own effect near g≈0.29 (Nickow, Oreopoulos & Quan, 2024, 89 randomized controlled trials).

one 2.0σ headline, four separate real numbers
Feedback-retest cycle
Tutoring's own share
Mastery learning (modern meta-analysis)
Tutoring (modern meta-analysis)
≈1.1σ of the original 2.0σ — the retest-until-threshold structure itself, no tutor present (von Hippel, 2024).
≈0.9σ — tutoring's own contribution once the feedback-retest cycle's share is subtracted out (von Hippel, 2024).
d≈0.52 on criterion tests, 103 studies, the largest mastery-learning meta-analysis (Kulik, Kulik & Bangert-Drowns, 1990).
g≈0.29, 89 RCTs, the most rigorous published tutoring estimate (Nickow, Oreopoulos & Quan, 2024).
All four numbers sit well below the original 2.0σ headline.

↑ Back to top

Test type: the hidden moderator

KEY-TERM

One variable swings these numbers more than any other: which test measured them.

The identical intervention can read two to six times larger on a test the researchers themselves built than on an independent, standardized test. Cohen, Kulik & Kulik (1982) found tutoring at 0.84σ on researcher-made tests against 0.27σ on standardized tests. Kulik, Kulik & Bangert-Drowns (1990) found mastery learning at 0.52σ on researcher-made criterion tests against roughly 0.08σ on standardized tests.

Test type is the single biggest moderator of how large an effect size looks. An effect size reported with no mention of which kind of test produced it says almost nothing on its own.

same intervention, same study family, 2-6x swing
Researcher-made narrow test
Independent standardized test
Large effect: tutoring 0.84σ; mastery learning 0.52σ (Cohen, Kulik & Kulik, 1982; Kulik, Kulik & Bangert-Drowns, 1990).
Much smaller effect: tutoring 0.27σ; mastery learning ≈0.08σ (Cohen, Kulik & Kulik, 1982; Kulik, Kulik & Bangert-Drowns, 1990).
The test type alone swings the same intervention's number 2-6x.

↑ Back to top

The percolation metaphor

CONCEPT

The second case comes from a different corner of the field entirely: network science.

A familiar metaphor — percolation — says a structure "connects up" once a random share of its pieces are switched on. Cross a critical fraction, and a single connected mass suddenly spans the whole thing.

The famous specific number attached to that story, p_c ≈ 0.5927, belongs to one exact geometry: site percolation on a two-dimensional square lattice. Leave that exact geometry, and the number carries no meaning at all. A different arrangement of connections has, in general, a different threshold — or no clean threshold whatsoever.

↑ Back to top

The geometry-free version

CONCEPT

The geometry-free version of this result is both more general and sharper.

Erdős & Rényi (1960) showed that in a random network, a single giant connected component appears suddenly once the average number of connections per point — the average degree, ⟨k⟩ — crosses one.

Below that crossing, a random structure almost surely fragments into many small, disconnected pieces. Above it, one piece suddenly dominates. This crossing, not any grid-specific number, is the real, general reason a thinly-covered space of interconnected things tends to fall apart into islands rather than staying whole.

↑ Back to top

Counting the pieces directly

CONCEPT

Now take the opposite case: one fixed, already fully known structure — not a random one whose connections are still being decided.

For that case, the right question isn't a probability threshold at all. It's a direct count of how many separate connected pieces the structure actually has. One piece means the whole thing is coherent. More than one piece means it has split into islands. A piece of size one means something is cut off entirely, connected to nothing.

That count is a standard, deterministic computation — breadth-first or depth-first search over the structure, not a statistical estimate. It applies exactly because the structure in question is already fixed and known, the opposite condition from the one a percolation threshold is built to handle.

MISCONCEPTION

Here's the misconception this sets up: assuming a percolation threshold, or its famous specific number, is the right tool to check whether a small, fixed, already fully known structure holds together.

It isn't. A percolation threshold is built for the opposite situation — a structure whose connections are still being randomly decided, where the question is a probability.

When the structure is already fixed and fully known, counting its connected pieces directly is both simpler and the actually correct check. It's a plain, deterministic computation — not a downgrade from a fancier-sounding, probability-flavored number that doesn't even apply to the case at hand.

↑ Back to top

What this chapter changes, and what it confirms

RECAP

Pull back, and both cases make the same point from different corners of the field.

A field's most quoted numbers are often an artifact of what was measured. Bloom's 2.0σ, decomposed and re-measured against far larger evidence, both shrinks and splits — a feedback-cycle share and a tutoring share, each near a quarter of the original headline.

Or the number is a metaphor borrowed from the wrong setting. A percolation threshold built for one specific random geometry stands in for a question about one small, fixed, already-known structure — a question a direct connected-components count answers more simply.

In both cases, the sturdier version isn't a compromise. Naming the test type behind an effect size, and counting connected pieces directly instead of reaching for a borrowed threshold, are both more accurate than the headline — and easier to actually check.

Education for One — Chapter 4: What the big claims are actually worth · projected from the LATTICE via prism_html.py · register: e41-graft

↑ Back to top