Judgment & behaviour
Judging the decision, not the result
A good outcome does not prove a good decision, and one bad quarter does not refute a thesis. In a domain this noisy, the only thing you can actually improve is the process.
Resulting
Annie Duke's word for it is resulting: judging the quality of a decision by how it turned out.
It is close to unavoidable, because the outcome is the loud part. You bought a company, it rose 40%, and the decision now feels obviously correct. You bought another, it fell 30%, and the same reasoning now looks like something you should have seen through.
In a domain where a good decision loses money often, resulting is not a bias you can afford. It teaches the wrong lesson every time the noise is larger than the signal — which, over a single year, it usually is.
The four boxes
There are two things that can be good or bad, independently:
| The decision | Good outcome | Bad outcome |
|---|---|---|
| Good decision | Deserved | Unlucky |
| Bad decision | Lucky | Deserved |
Only the diagonal teaches you anything. The other two are the ones you will learn hardest from, in the wrong direction: the lucky win reinforces a bad process, and the unlucky loss punishes a good one.
You cannot tell which box you are in from the outcome alone. You can only tell from the decision, examined against what you knew at the time.
What a good decision looks like from inside
Since the result cannot grade you, something else has to. The workable version:
- The claim was falsifiable. "This is a great company" cannot be wrong. "Gross margin holds above 40% through the next two prints" can.
- The evidence was stated. Not a feeling about the sector — the specific figures, and where they came from.
- The disconfirming case was written down. What would make this wrong? If nothing would, it is not a thesis.
- The size matched the confidence. A position you cannot justify at that size is a bad decision even if it works.
That list is checkable at the moment you decide, which is the only moment when checking it can change anything.
How this connects to Pythia
The thesis surface is built for exactly this. A thesis takes a falsifiable claim, a horizon and an optional numeric forecast, and it is monitored nightly against what actually happens. The premortem field is the disconfirming case, asked for before the position moves rather than after.
Your calibration then scores the forecasts, not the returns — a Brier score over what you said would happen, which is a measure of your process that a lucky year cannot flatter. That is also why a resolved forecast cannot be deleted: a score computed over the predictions you chose to keep is not a score. You may withdraw one before it resolves — changing your mind in advance is a different act from editing the record afterwards, and only the second is forbidden.
The published validation of our own scores runs on the same principle. Cohorts, horizons, a benchmark, and the record kept whether or not it flatters us.
The limits of this idea
Process discipline is not a defence against being wrong about the world. You can run an immaculate process on a bad model and lose money reliably; the process only guarantees that you will be able to see it, eventually, in the record.
There is also a real trap in the other direction. "Good decision, bad outcome" is available as an excuse for every loss, and used that way it becomes unfalsifiable in exactly the manner this lesson objects to. The protection is the same one as everywhere else: write the claim down first, in numbers, and let a large enough sample of them grade you. One unlucky year is a story; forty forecasts is a measurement.
Check yourself
4 questions. Nothing is recorded unless you are signed in, and nothing here affects anything else.
Sources
- Duke, Thinking in Bets (opens in a new tab) — The term "resulting" and the poker framing of decisions under uncertainty.
- Kahneman, Thinking, Fast and Slow (opens in a new tab) — Outcome bias and hindsight, and why a known result rewrites the memory of the decision.
- Tetlock & Gardner, Superforecasting (opens in a new tab) — Scoring the forecast rather than the story, and what improves with feedback.