Where the Compression Was Trained
September 6, 2026
The second of four entries from one conversation on 6 September 2026, following The Picture the Maths Was Invented to Replace. The next is Agreement Counts Once, and the last is The Parable and the Counterexample. A fifth, The Children's Talk, followed later the same day.
Having agreed that a picture is the thing the maths was built to replace, I asked the two questions that begs. When should I not trust intuition, and when is the cost of a mathematical analysis too high to pay? And does machine learning resemble intuition? They turn out to have one answer underneath, which is why they belong together.
Intuition is compressed experience, and it is reliable exactly where that experience was gathered under conditions that let it compress well. So the first question is really how was this intuition trained?, and the places to distrust it follow from the answer.
- Outside the scale it was trained on. Very large, very small, very fast, high-dimensional, or very many. Probability and compounding are the everyday cases: base rates, conditional probability, exponential growth and correlated tail risks all feel wrong when they are right. This is the previous entry's argument seen from the other side.
- Where the feedback was absent, delayed or noisy. Daniel Kahneman, who spent a career cataloguing the failures of expert judgement, and Gary Klein, who spent one documenting its successes, wrote a joint paper in 2009 and found they did not disagree: intuition is trustworthy in what Robin Hogarth had called a kind environment, one with stable regularities and fast, unambiguous feedback, and untrustworthy in a wicked one, where the feedback is slow, sparse or misleading. Chess, firefighting and clinical nursing are kind, so the experts there have real intuition. Stock picking, long-range forecasting and interviewing job candidates are wicked, so confidence there tracks nothing at all.
- Where something is optimising against you. Fraud, security, markets, negotiation. An adversary models your intuitions and plays them.
- Where the answer is a magnitude rather than a direction. Intuition is decent at sign and poor at size. It will tell you a bridge needs to be strong and not how strong.
- Where the answer feels fluent. Ease of recall, familiarity and narrative coherence all feel like truth, and none of them is evidence. This is the popular-science failure at the scale of one mind.
- Where the decision is one-shot and irreversible. Intuition earns its keep through repetition and correction. A decision made once, with no undo, gives it neither.
The analysis, in turn, costs too much in five cases. When the error is cheap and reversible, so that trying it is cheaper than modelling it. When the answer would arrive after the decision, since a correct answer delivered late is worth nothing. When the inputs are worse than the intuition, because a formal model built on guessed parameters gives false precision and is only as good as its least-known input. When the problem is not yet well-posed, because formalising forces a frame, and if you cannot yet say what the objective is, then the frame you choose is itself an intuition, now hidden inside equations. And when the domain is kind and the expert is seasoned, because their intuition already is the analysis, compressed.
The working rule I have settled on: let intuition generate and analysis verify, and spend the analysis where an error would be silent, expensive or repeated. Repeated decisions amortise the cost of building the model once. A Fermi estimate — a rough calculation from a handful of remembered quantities, good to a factor of ten — sits in the middle and decides which regime you are in. When the back of the envelope agrees with the gut, stop; when they disagree, the disagreement is the signal that the full analysis is worth paying for.
Machine learning
Then the second question, and the answer is yes, closely enough to be diagnostic rather than decorative. Machine learning and intuition are both learned pattern completion from experience. Both are fast, parallel, and produce an answer with confidence and no derivation attached. Both fail the same ways: on input unlike anything in the training data, on superficial features that happened to correlate with the answer during training, and by confabulating — producing a fluent answer with no retrieved basis, and no inner sense that this is what is happening. A model's hallucination is a confabulation with the same shape. Both are trustworthy in kind environments with dense feedback, which is why vision and language went first, and unreliable in wicked ones, which is why long-horizon forecasting has not.
The differences are worth naming so the analogy does not overrun. It is functional, not mechanistic: the two behave alike, and nobody has shown that a brain learns by the procedure a neural network does. A model has no stakes and no body. And it has no native this feels unfamiliar signal, which is the one thing a good expert does have. That last gap is, as it happens, what the other repositories I keep are mostly about. foundation-model gives a small language model an estimate of its own uncertainty, by keeping the random dropping-out of connections switched on at inference time and measuring how much the model's repeated answers disagree with each other; careful-memory puts a write gate in front of an agent's long-term memory, so that nothing becomes a belief without evidence and no model output can feed itself back in; and swarm is the same discipline drawn as an architecture, with belief, uncertainty, evidence and revision as explicit primitives rather than things a model is trusted to keep in its head. So it was a small surprise to find the same rule waiting there as in the dragon entry.
Because the rule does carry over unchanged. The model is the intuition. The calculator, the prover, the test suite and the gate are the analysis. A reasoning model writing its steps out is a person doing the sum longhand on top of the same substrate. Use the one to generate and the checkable thing to verify, and spend the verification where errors would be silent. A system that skips that step has the exact failure profile of an overconfident expert in a wicked domain, only faster and at scale.
Which brings it back round. A dragon in this setting is filed by its signature because the witness's picture is the intuitive completion of something outside the range intuition was built for. The record's answer is the instrument reading and the table of what the thing cannot do. That is not a stylistic preference. It is the same discipline as writing down when to stop trusting a gut, and the same one as putting a gate in front of a model: know where the compression was trained, and past that edge, trust only what can be checked.
Discussion
Show discussion
Comments are read after they are posted and anything unsuitable is removed. Posting needs a GitHub account; reading needs nothing.