
What It Is
An explanation can be internally consistent without being true. It can fit many observations into a small model, contain no obvious contradictions, and feel complete. Those qualities belong to the explanation. Evidence comes from checking it against the world: "coherence is a property of the story; evidence is a property of the world."
When external evidence is absent, the mind treats coherence as evidence because it has no other signal available. Its generative model produces confidence from internal features such as compression, explanatory reach, and consistency, rather than the model's observed hit rate. Confidence regulates the system: it calms tension and ends a search. A belief that achieves that relief gets accepted whether or not it was tested.
Reality contact describes the practice of obtaining feedback. Reality contact metabolism describes the physiological account of contact as a nutrient and anxiety as a deficiency signal. The additional requirement is a promotion policy: a rule for what a hypothesis must pass before it can be treated as a belief. Without one, whatever feels coherent gets promoted, leaving a sophisticated mind full of confident hallucinations.
A closer look
Two different sources of confidence
Internal consistency helps form a hypothesis. An observation or test supplies a different kind of support.
Read this diagram
Compare The story fits together; The world supplies evidence.
The Origin Scene: Catching the Catch
Will noticed the problem when a confident thought appeared:
"There are no papers about agent convergence."
He then questioned where it came from:
"How do I know this? I haven't checked arxiv. This is a hallucination coming from lack of evidence. I have not sampled."
The claim was specific, declarative, and ready to support further reasoning. It felt no different from a checked fact. Asking about its provenance revealed that it had been generated from a guess, not retrieved from an observation. The same provenance question can evaluate markets, cached conclusions, and agent output; here it determines whether a thought has earned belief status.
Watching language models helped Will recognize the process in himself. Thousands of plausible but ungrounded answers showed how cheaply a generator manufactures fluency. Once he could recognize the AI's smooth wrong answers, he could recognize his own. The two shared the same architecture and lacked an internal marker distinguishing sampled knowledge from confident invention.
That observation made the thought available for inspection as a process: a proposed branch, a confidence signal, and a reward loop. In predictive coding, the brain is a generative model, and generative models hallucinate by design.
Confidence Is Regulatory, Not Epistemic
The confidence signal was built to close open loops, calm arousal, and free working memory for the next problem. It was never built to track truth. Feeling done reports that regulatory state, even when the explanation has not been checked:
"A model can feel real without being correct — coherence is a property of the story, evidence is a property of the world."
The feeling is difficult to distrust because it successfully resolves tension. That success does not verify the belief that produced it.
| Signal | What it actually tracks | What you read it as |
|---|---|---|
| "This makes sense" | The explanation is internally consistent and legible | The explanation is true |
| "This feels done" | The search loop has terminated; tension resolved | The work is complete |
| "This is elegant" | High compression, few moving parts | High probability of being right |
| "I could explain this to anyone" | The story is fluent | The mechanism is mastered |
| "Obviously X" | No counterexample is currently loaded in working memory | No counterexample exists |
Each substitutes a property of the representation for a property of the world. The substitution happens silently whenever the external signal is missing. For an intelligent person working in simulation, it is usually missing. Better modeling produces more coherent stories, stronger confidence, and less felt need to check. The trap tightens with IQ.
Will described the grandiose version in the metabolism article: "when I feel like I'm a genius, it's usually because of lack of reality contact." Maximum coherence with zero sampling produces that feeling. The model has encountered no contradiction because nothing outside it has been allowed to supply one.
The Promotion Policy
Keep generating bold, specific models. They are valuable hypotheses, and suppressing the ability to produce them would discard that value. Change the rule that promotes them instead.
Before treating a hypothesis as a belief, ask four questions:
- What is the evidence surface? Identify the part of the world the claim concerns: an arxiv listing, a user, a scale, or a market. If there is no nameable place to check, the claim is not about the world.
- Have I sampled? Determine whether the claim came from contact or was generated from prior expectations. Generated claims are usable as hypotheses.
- What would falsify this? Name a possible finding that would contradict it. Without one, it is a mood expressed as a claim rather than a belief.
- What is the cheapest reality check? Identify the smallest useful search, message, or measurement, then perform it.
Generation and promotion become separate stages. Produce models at full rate, but use an untested one only to design a probe. It cannot yet serve as a foundation for further conclusions. This applies selection-over-design to beliefs: generate possibilities, test them, and promote those that survive contact.
"If I haven't checked... nothing exists to push back, and the internal model becomes sovereign by default."
An unchecked model becomes the standard against which later thoughts are evaluated. Its errors then pass into downstream reasoning as axioms. This is epistemic contamination: one unearned belief contaminates every inference built on it. The problem remains invisible because those inferences agree with the assumption that generated them.
A check also gives the search a boundary. Coherence can always be increased, so refining a model can continue indefinitely, with each story generating more sub-stories. “What is the cheapest reality check?” turns that open task into one bounded sample. Bounded search supplies the same principle: a search without a budget defers selection, while a belief without a falsification test defers sampling.
What Coherence Is Actually For
Coherence remains useful before promotion and after evidence has been collected:
| Role | Coherence used as | Legitimate? | Why |
|---|---|---|---|
| Hypothesis quality | Ranking which unsampled models are worth testing first | Yes | A coherent, compressive, mechanism-shaped hypothesis has a better prior than an incoherent one |
| Probe design | Deriving what the model predicts, so you know what to check | Yes | You cannot falsify a story that makes no commitments; coherence makes the commitments legible |
| Compression | Packaging already-sampled knowledge for transfer and reuse | Yes | Post-evidence coherence is what a good map is made of |
| Promotion | Upgrading the hypothesis to a belief | No | This is the swap — an internal property standing in for an external one |
| Reassurance | Ending the discomfort of not-knowing | No | This is the regulatory function running the epistemics |
A coherent model makes a sharper prior, states predictions that can be checked, and packages tested knowledge for reuse. These abilities make someone a better searcher. In Bayesian terms, coherence initializes the estimate; only contact updates it. Route coherence into ranking hypotheses and designing probes, while keeping it out of the promotion decision.
Dense feedback enforces that separation automatically. Code runs or fails. A barbell rises or does not. The result supplies a check without a separate effort to arrange one. The problem becomes consequential where feedback is sparse, delayed, or avoidable: strategy, self-assessment, markets not entered, and papers not searched. Strong simulators prefer precisely these domains. An explicit gate supplies the check that the environment does not impose.
The Three Levels: Territory, Map, Index
The same mistake can promote a claim about competence: “I know this subject.” Distinguish three things someone might possess:
| Level | What it is | What possessing it feels like | What it actually gives you |
|---|---|---|---|
| Territory | The actual work — code run, problems drilled, kinks hit | Ordinary; full of friction and specifics | Operational capability |
| Map | A systematic representation — the textbook, the course, the framework | Structured understanding | Navigation, if minted from contact |
| Index of the map | The list of names — concepts, terms, titles | Fluent parity with experts | Pattern-matching without mechanism |
Fluency with the names can be mistaken for understanding the representation, and understanding the representation can be mistaken for practical contact. Will described his learning error:
"I thought the map was the territory, and that learning meant memorizing the map rather than using the map to get acquainted with the territory. Worse — I thought reading the index of the map was understanding the map."
He traced it back to an early attempt to learn about compilers:
"I was a nerdy asian kid, glasses... genuinely curious about compilers enough as a 6th grader that I found the resources — tried to read the dragon book, COULDN'T lmao, but still pretended."
He had the curiosity and intelligence to find the material, then spent that intelligence maintaining the appearance of understanding it. The pattern continued for two decades: feeling equal to practicing engineers "cuz I knew the name of a concept and could pattern match," or watching OpenCourseWare lectures and counting that as having taken the course.
Learning the index is cheap and supplies enough fluency to pass in conversation. Social approval then confirms the supposed competence that actual performance would have disproved. That feedback keeps the mistake intact.
The missing knowledge is sub-verbal: "Intelligence / knowledge is embodied, not in the traces." A practitioner has learned "the tiny nuances about an algorithm that are too trivial to write down": which data structures resist the approach, where a method fails, and what particular error messages mean. Those details exist only through contact. A map records what is worth writing down and leaves out these small operational nuances.
Zhuangzi's wheelwright describes the same limitation through a feel in the hands that cannot be transmitted. Knowledge that can be indexed carries the least evidence of mastery. Skill acquisition and pedagogical magnification develop the requirement: understanding is real only at the resolution where you can act.
The Corrective Filter: Demand the Mechanism
Apply the test when evaluating other people as well as yourself. Looking at the Heads of Education at two competitors on LinkedIn, Will noticed himself judging their schools, titles, and appearance. He was evaluating metadata rather than their work:
"If I wanted to be better than them I need a mechanistic interpretation of what makes their videos suck and a working theory of what makes a video good — informed by reality rather than arbitrary 'good'."
Credentials index someone else's map. The only discriminating question is whether that person can explain and execute the causal chain.
Apply the same question when your own vocabulary becomes fluent. Until demonstrated otherwise, easy conversation about a domain shows fluency with its index. A word can be a handle or a blindfold: knowing the term “backpropagation” feels like knowing the mechanism, which can suppress the check.
Will's earliest recorded version was six-year-old reasoning: "to grow taller I could just play basketball". He mistook the name of a correlated activity for a mechanism that would produce the outcome. Adult versions can repeat the error with a more elaborate explanation.
"Reading the index is not reading the map, and reading the map is not walking the ground."
Comprehension Is a Read; Mastery Is a Write
Following an explanation shows that you can read it. It does not write durable capability. Comprehension is a read operation; mastery is a write operation. The feeling of understanding arrives before the ability to perform:
"This laxness with feeling mastery is unnecessary — usually I'll stop when my conscious mind is able to [follow it]."
The conscious gauge reports completion at maybe 30% of the actual work. It measures whether the explanation is legible, not whether the capability is durable. The gap appears later during an interview or live build, when the feeling of fraudulence finally supplies an honest but expensive signal.
"It took me till I'm 29 to learn that you can't just glance over and read and have it just 'make sense.'"
Use an execution test instead: "basically can I execute the thing, not just know the entire causal chain once." Take the code apart and rebuild it. Reproduce the explanation cold, without looking at the source. Closed-book execution is the only read-out that fluency cannot forge. Comprehension supplies a hypothesis of competence; execution supplies the sample. An autodidact framework needs this stopping condition to avoid becoming repeated map consumption.
| Stop condition | Operation type | What it verifies | Failure mode |
|---|---|---|---|
| "It makes sense" | Read | The source is legible | Closes the tab at 30% |
| "I can summarize it" | Read with compression | The index is loaded | Fluent name-dropping |
| "I could do it with the source open" | Assisted write | The map navigates | Collapses without scaffold |
| "I can execute it cold" | Write | The territory is walked | None — this is the gate |
Arrange the practice so that execution is required:
"You have to join a structure where doing the 'hard' things is thermodynamically optimal."
Re-deriving a proof from memory always requires more energy than rereading it. Willpower loses against that difference; the Boltzmann distribution does not change with intentions. A job, cohort, deadline, or person waiting for an artifact changes the cost of avoiding practice. Execution becomes the path of least resistance, and merely feeling that you understand no longer satisfies the requirement. Container design and forcing functions make not practicing the expensive option.
AI increases the temptation to stop at comprehension. It can produce perfect explanations at any depth and in any style indefinitely. If “makes sense” is your stopping condition, that is what you keep requesting. The tool could provide unlimited drills, but it defaults to the explanations you ask for. Each session can end with complete felt comprehension and no executed practice. Request the drill deliberately. AI as Accelerator multiplies whichever loop it enters, including a loop made entirely of reading.
The Dual-System Point
Humans and generative models fail identically here. Both generate plausible, confident, mechanism-shaped claims without an internal marker that distinguishes sampled knowledge from invention. Both sound most convincing while elaborating a coherent story, regardless of whether it is true.
For AI systems, this requires verification against external results. A verifier or golden set of known-correct cases implements the promotion policy. Asking the generator to be more careful does not establish groundedness. Taste compilation describes how judgment becomes such tests. The tests are necessary because fluency is cheap and supplies no evidence by itself.
For someone thinking with AI, there are now two generators: one produces coherent but ungrounded models at biological speed, and the other at machine speed in a form tailored to be legible to that person. They can confirm one another without either having checked the claim. Each mistakes the other's fluency for corroboration.
That makes an explicit promotion policy more necessary than ever. The gate questions are the only protection against a workspace full of mutually consistent hallucinations supplied by both the person and the model.
Failure Modes
| Failure mode | Signature | The patch |
|---|---|---|
| Confident confabulation | Specific factual claim, zero provenance ("there are no papers on X") | Provenance audit: generated or retrieved? Run the cheapest check |
| Sovereign model | Whole plans built on an unchecked premise; disagreement feels absurd | Name the premise; name what would falsify it; sample before building further |
| Index masquerade | Fluent vocabulary, zero executed reps in the domain | Demand the mechanism of yourself; count contact-hours, not concepts named |
| Credential filtering | Judging people/artifacts by titles, schools, cosmetics | Demand the mechanism of them; evaluate the causal chain, not the metadata |
| Comprehension stop | Tab closed at "makes sense"; gaps surface at performance time | Closed-book execution as the only stop condition |
| Coherence squared | Long AI sessions elaborating an unsampled model; everything agrees | Deliverable-gated sessions; sample between generations; abstain when polluted |
| Generator suppression | Distrusting all bold hypotheses to avoid false confidence | Wrong patch — gate the promotion, don't dampen the generator |
Generating fewer bold hypotheses would discard useful signal without repairing automatic promotion. Preserve strong models as candidates and change what they must pass before being relied on.
Running the Gate: Practical Implementation
Use the checks at the moments confidence makes them seem unnecessary.
Trigger 1: A confident factual claim appears mid-thought
Claims such as “there is no X,” “everyone does Y,” and “the market wants Z” can arrive already feeling established.
- Audit provenance. Name the search, conversation, or data from which the claim came. If no sample can be named, it was generated.
- Check before using it as a premise. For the claim about agent-convergence papers, one arxiv search takes thirty seconds. After a claim becomes a premise, it stops being visible as a claim, so the order matters.
- Log the catch. Each caught confabulation measures a base rate higher than the feeling reports. Keeping that record sustains the reason to perform the audit.
Trigger 2: A domain starts feeling mastered
Fluent vocabulary, connected concepts, and a feeling of parity with experts should trigger a performance check.
- Count contact-hours: exercises completed, difficulties encountered, and procedures executed. A near-zero count means the fluency belongs to the index.
- Attempt a closed-book write. Rebuild the thing, derive the argument, or run the procedure from memory. Compare that result with the feeling of competence.
- If the domain matters, enter a cohort, job, deadline, or relationship with someone waiting for the artifact. Make practice thermodynamically optimal through that structure instead of relying on willpower.
Trigger 3: A long generation session with no samples in it
Hours of elaboration, alone or with AI, can make every part of a model increasingly consistent with every other part.
- Count the external samples: searches, conversations, and executed tests. Zero samples with rising confidence is the alarm condition.
- Identify the one or two untested premises supporting the model. Write each as a hypothesis and give it a falsification test.
- Finish with a probe specification. An unsampled session cannot produce a belief; it can specify the cheapest experiment. Run that probe before the next elaboration session.
These checks replace confidence, fluency, and completion feelings with counts of samples, contact-hours, and executed writes. Coherence can forge feelings but cannot forge those counts. Tracking puts the record outside the mind whose internal record has already been altered by simulation.
Integration with the Mechanistic Framework
Connection to Reality Contact
Contact supplies samples, the only basis on which a hypothesis can earn belief status. Sampling without a promotion policy still allows coherence to decide what gets believed. A policy without sampling promotes nothing. Both are required.
Connection to Reality Contact Metabolism
The metabolism account describes anxiety, grandiosity, and depression under contact deprivation. Deprivation also misleads: with coherence as the only remaining signal, an isolated model grows more confident as it grows more coherent. The genius feeling reports the regulatory effect of accepting unchecked beliefs.
Connection to EV Sensor Calibration
Virtual EV applies the same substitution to value. A simulated payoff becomes felt certainty without lived samples. Exposure alone reprograms those sensors. The gate asks what an outcome is worth instead of whether a proposition is true.
Connection to Predictive Coding
A generative brain minimizes prediction error. Without incoming samples, there is no error to minimize; its model elaborates unopposed and its precision weighting defaults to trusting itself. Confidence without samples follows from running that architecture without feedback.
Connection to Pattern Matching
Index knowledge can recognize, classify, and converse without being able to execute a mechanism. Pattern matching becomes a legitimate shortcut once the mechanism is embodied. Before that, recognition can be mistaken for competence.
Connection to Taste Compilation and Bounded Search
Taste compilation turns the promotion policy into golden sets, verifiers, and mechanical specifications. Bounded search limits indefinite elaboration. Falsification tests and the cheapest available check supply stopping conditions for an otherwise unbounded search.
See Also
- Reality Contact — obtaining feedback and distinguishing simulation from results
- Reality Contact Metabolism — contact as nutrient and anxiety as a deficiency alarm
- EV Sensor Calibration — untested confidence about value
- Predictive Coding — the generative process behind confabulation
- The Matrix — mistaking the coherent rendering for what generates it
- Epistemic Contamination — downstream consequences of an unearned belief
- Autodidact Framework — requiring execution before declaring learning complete
- Skill Acquisition — the sub-verbal knowledge acquired through practice
- Pedagogical Magnification — understanding at the resolution needed to act
- Pattern Matching — recognition without an executable mechanism
- Language Framework — vocabulary as both handle and blindfold
- Clarity Bear — the adjacent risk of mistaking a clear vision for a valid one
- Taste Compilation — implementing the checks in machinery
- Bounded Search — terminating elaboration with a test
- Selection over Design — generating freely and promoting after selection
- Provenance — the generating path determines evidential weight; this audit is its founding instance
- The Simulated Other — mistaking a rendered person for a checked understanding of them