statistical-power
The chance an experiment catches a real effect if one is there — the net's mesh size: too coarse and the fish swims through, too fine and you drown in effort.
Power is the question that decides whether a design is worth running at all. The free-choice gap diagnostic wing asks whether the two-unrewarded-tasks gap (predicted d ≈ 0.20–0.35) is large enough to detect at a feasible class size (minimum-class-size), and the teachability validation asks whether 10–12 concepts × 15–20 learners gives enough signal to test a ranked calibration (cheapest-teachability-validation). The trap is symmetric: too little power and you miss a real effect (a false negative), too much and you detect an effect too small to matter (a trivially true positive). The honest path the castle's design rooms follow is to run the simplest version first to estimate the effect size, then power the full study to the measured gap — because guessing the effect size from a benchmark built for a different question is the wrong comparison (two-task-effect-size).
Links
effect-size
How big a difference really is — not whether it exists, but whether it is large…
WORD · brickwithin-subject
A within-subject design is one where the same person does every condition — so t…
WORD · brickfree-choice
A way to measure intrinsic motivation: after the task ends and no one is watchin…
ROOM · wallIf the class-level gap difference diagnoses the task but the free-choice measure is notoriously noisy, what is the minimum class size that reaches significance — and does the informational reveal's gap-change have enough effect size to clear the noise bar at that class size?
The stethoscope pressed to a hundred chests hears the fever the single pulse drowned in — but only if the fever is louder than the ward's own murmur.
ROOM · wallWhat is the cheapest design that would validate a student model's teachability score against human learning — and how many concepts are needed before the correlation is signal, not noise?
The mannequin wore every coat to perfection; the question is how many children must try them on before the tailor's rankings can be trusted.
ROOM · wallIf the gap difference between two unrewarded tasks of different value may be smaller than the reward-undermining effect (d = .28–.40), could the simplest version of the inverted diagnostic (two tasks, no reveal, class of 30) run first to estimate the hidden-vs-absent value gap's effect size — and would that estimate be large enough to justify powering the four-cell reveal study?
Before you build the telescope, hold the ruler to the star — if the light is too faint, no glass will catch it.