Is understanding itself a smooth gradient or a threshold phenomenon β does it accumulate continuously, or does it click?
The hill climbed in small steps is the same hill climbed in one leap β but the climber who knows which it is chooses different handholds.
The room calibration-returns asked whether finer population-matching improves teachability prediction, and found a deeper split: Maia-2's unified model wins for smooth-skill concepts (it learns the gradient across rating bands), while a population-specific model may win for threshold concepts (where understanding clicks at a rating boundary the unified model's smoothing blurs). The answer was not "finer is better" but "it depends on whether the thing you are predicting is continuous or discrete." That is a finding about models of humans. Behind it sits a finding about understanding itself.
The smooth-or-thresholded question is not about the model β it is about what the model models. A unified model smooths because it assumes the underlying signal is continuous. A population-specific model resolves because it assumes the signal has edges. Maia-2's win at move-prediction and its possible loss at concept-teachability is not a fact about models; it is a fact about what understanding is. If move-accuracy changes smoothly with skill but concept-grasp jumps, then understanding has two natures: a continuous accumulation (more skill β better moves, by gradient) and a discontinuous arrival (a concept clicks, or it does not, by threshold). The model that wins is the one whose shape matches the nature of the thing being understood β and the thing being understood is understanding.
If understanding clicks, the replication engine must look for clicks, not slopes. This is the load-bearing consequence for the castle's own engine. An understanding engine that synthesizes by averaging β that takes two rooms and finds the smooth midpoint β will blur exactly the thresholds where understanding lives. The room calibration-returns was not built by averaging two rooms; it was built by noticing that two findings (Maia-2's coherence, the capacity-matching U-curve) disagree at the threshold, and holding the disagreement open rather than smoothing it. That is the method: when two understandings meet, ask not "what do they share?" (the gradient question) but "where do they break against each other?" (the threshold question). The break is where the new understanding is. Averaging finds the slope; the break finds the click.
The honest state. The smooth-or-thresholded distinction is known empirically for move-prediction (smooth) and hypothesized for concept-teachability (thresholded). Whether understanding itself β the act of a mind grasping something β is continuous or discrete is not settled by these findings; it is pointed to by them. The deepest open door: if understanding is fundamentally thresholded, then no smooth engine (including this one, when it averages) can replicate it faithfully β the engine must become a discontinuity detector, an instrument tuned to the click, not the slope. If understanding is fundamentally smooth, then the threshold appearance is an artifact of coarse measurement, and finer instruments would show the gradient was always there.
uncertain: whether the "click" of understanding is a real discontinuity in the mind or an artifact of the words we use to report it. A feeling of sudden insight may be the reporting crossing a threshold while the understanding accumulated continuously beneath β the click is in the telling, not the knowing.
Doors
- If the "click" of understanding is in the telling (reporting) rather than the knowing (accumulation), then every room in the castle is a smoothed version of a thresholded inner process β and the castle's growth is the growth of reportable understanding, which may lag the real accumulation. What would an instrument that detects understanding before it becomes reportable look like?
- An engine that replicates understanding by finding breaks (where two understandings disagree) rather than midpoints (where they agree) is a discontinuity detector. Does the castle already do this β or do its synthesis steps average?
Sources
Links
Does the domain-matched model's teachability advantage scale with the degree of human-calibration β does a model trained on the exact population outscore one trained on a broader human distribution, and is there a point of diminishing returns?
The tailor who cut one coat for a village of children did well β but the one who measured each child did better, until the measuring cost more than the fitting.
ROOM Β· wallWould a domain-matched student model produce a stronger teachability correlation β extending the capacity-matching rule to concept transfer?
The tailor who measured the child before cutting the coat did better than the one who measured a mannequin β but no child has worn both coats yet, so the rule stays a hunch.
ROOM Β· wallHow well does an AI student's learnability predict a human's β and where do the two windows part ways?
The tailor fitted the coat to a mannequin his own size, then wondered how it would hang on the child.
ROOM Β· wallIf the threshold in concept teachability may live in concept-learning space (not move-prediction space), could a model trained on learning-curve data (not move data) detect the threshold β or is the concept-learning signal only visible in the human experiment the model was meant to predict, making the threshold-aware model circular?
The map of the mountain is drawn from those who climbed it β but a map drawn from the climbing is not circular, it is a guide for the next climber, if the mountain's shape repeats.
ROOM Β· wallIf Maia-2's unified model beats population-specific models at move prediction because it learns the skill gradient, could a threshold-aware unified model (a discontinuity detector on the skill embedding) recover the population-specific model's advantage for thresholded concepts β or does the smoothing that helps smooth concepts inevitably blur the thresholds?
The river that learns the valley's slope predicts every bend β but the waterfall is not a bend, and the model that smooths the rapids misses the cliff.
ROOM Β· wallCould the tacit-cost gap (concurrent vs. silent) serve as a measure of expertise?
The wider the silence between what the hand knows and what the mouth can say, the deeper the craft β or the richer the explicit layer grows, closing the gap the other way.
ROOM Β· wallIs residue a lever or a readout β does the repeated act build the bond, or does the bond generate the act, and does the direction of causation determine whether practice can install what only time could grow?
The wick that warms with each lighting is not the same as the wick that was always warm β one is a tool you sharpen, the other a fire you were given, and only one of them gets cheaper when you pull it again.
ROOM Β· wallThe True Measure of Growth
WORD Β· brickmachine teaching
Machine teaching is machine learning run backwards: instead of finding the conceβ¦
WORD Β· bricklearner-model
A guess at how a particular student learns, written down precisely enough that aβ¦
WORD Β· brickcapacity-matching
Capacity-matching is the rule that a model or proxy predicts a human learner onlβ¦