Accountability
We keep score.
Every day we record a verdict for the whole universe — a PAS (Academic Score), a Guru Pyramid tier, a best-match guru. This page measures how those cohorts actually performed against the S&P 500 afterward. It is computed mechanically, never hand-picked, and it is a forward, out-of-sample record — not a backtest.
Scorecard vintage September 11, 2026 · cohorts May 16, 2026 → Aug 12, 2026 (71 verdict days) · live since May 16, 2026
Top decile vs SPY · 90d
+1.0%
n = 7,155
Top-decile win rate vs SPY
48%
share of names beating SPY
Decile 10 − decile 1 spread
-74.3%
the score's ranking power
Does the score rank the future?
Pooled excess return vs SPY per Pythia decile at the 90-day horizon. A working score slopes upward: low deciles underperform, high deciles outperform.
Current read: the top decile is underperforming the bottom decile by 74.3 pp — the early-window sample is not yet ranking forward returns; we publish it anyway. Thin-coverage names can inflate scores until the coverage gate ships.
Cohort by cohort
Each point is one verdict day's 90-day forward return — the top decile, the bottom decile, and SPY over the identical windows. No pooling, no smoothing.
The full scorecard
By PAS decile
Decile of each cohort's own PAS distribution (10 = highest).
n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature
| Bucket | Avg return | SPY | Excess (median) | Win rate | P(beat SPY) | n, names pooled | Cohorts, verdict days |
|---|---|---|---|---|---|---|---|
| Decile 10 (highest) | +2.0% | +1.3% | -0.8%mean +0.8% | 46.6% | 49.1%9% pooled | 35,140 | 63 |
| Decile 9 | +3.5% | +1.2% | -0.2%mean +2.3% | 47.7% | 51.8%9% pooled | 34,770 | 63 |
| Decile 8 | +2.5% | +1.3% | -0.3%mean +1.2% | 47.0% | 52.6%9% pooled | 37,912 | 63 |
| Decile 7 | +6.8% | +1.0% | +0.3%mean +5.8% | 49.6% | 52.8%9% pooled | 27,103 | 63 |
| Decile 6 | +4.4% | +1.4% | -0.5%mean +3.0% | 45.8% | 51.8%9% pooled | 41,796 | 63 |
| Decile 5 | +6.2% | +1.3% | -0.7%mean +4.9% | 46.2% | 50.1%9% pooled | 37,563 | 63 |
| Decile 4 | +15.6% | +1.2% | -0.6%mean +14.4% | 47.1% | 49.5%9% pooled | 32,389 | 63 |
| Decile 3 | +11.0% | +1.1% | -0.7%mean +9.9% | 46.9% | 47.8%9% pooled | 34,071 | 63 |
| Decile 2 | +30.9% | +1.3% | -1.4%mean +29.6% | 44.6% | 44.2%9% pooled | 34,283 | 63 |
| Decile 1 (lowest) | +37.3% | +1.2% | -3.0%mean +36.1% | 41.9% | 39.7%9% pooled | 34,226 | 63 |
By pyramid tier
Canonical pass-count band at the time of the verdict.
n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature
| Bucket | Avg return | SPY | Excess (median) | Win rate | P(beat SPY) | n, names pooled | Cohorts, verdict days |
|---|---|---|---|---|---|---|---|
| DEEP VALUE | +3.1% | +1.6% | -0.3%mean +1.5% | 49.0% | 54.0%15% pooled | 3,457 | 61 |
| QUALITY VALUE | +3.1% | +1.3% | +0.2%mean +1.8% | 49.8% | 55.9%14% pooled | 37,100 | 63 |
| FAIR VALUE | +3.4% | +1.1% | +0.1%mean +2.2% | 49.5% | 53.3%14% pooled | 112,110 | 63 |
| OVERVALUED | +17.4% | +1.0% | -1.0%mean +16.4% | 45.2% | 46.6%14% pooled | 288,597 | 63 |
By best-match guru
The guru archetype each company most resembled that day.
n = names pooled · cohorts = verdict days · P(beat SPY) = per-cohort win rates shrunk toward the family average — estimated in-sample and graded as new cohorts mature
| Bucket | Avg return | SPY | Excess (median) | Win rate | P(beat SPY) | n, names pooled | Cohorts, verdict days |
|---|---|---|---|---|---|---|---|
| Marks Cycle | +19.1% | +1.0% | -0.4%mean +18.1% | 48.0% | 48.1%23% pooled | 115,231 | 63 |
| Pabrai Value | +20.9% | +1.2% | -1.1%mean +19.6% | 44.3% | 46.2%23% pooled | 68,619 | 63 |
| Buffett Moat | +12.4% | +0.7% | -0.6%mean +11.8% | 46.7% | 46.1%23% pooled | 55,619 | 63 |
| Safety First | +2.3% | +0.8% | -0.5%mean +1.5% | 46.8% | 50.0%23% pooled | 36,092 | 63 |
| Greenblatt Magic | +4.0% | +0.9% | +0.7%mean +3.1% | 52.7% | 52.4%23% pooled | 29,961 | 63 |
| Drucker Efficiency | +0.9% | +1.7% | -1.6%mean -0.8% | 39.9% | 47.1%29% pooled | 24,055 | 46 |
| CAPEX Efficiency | +4.8% | +2.2% | -1.8%mean +2.6% | 41.0% | 45.0%49% pooled | 13,237 | 19 |
| Thorndike Outsiders | +2.9% | +2.5% | -1.5%mean +0.3% | 43.0% | 43.9%28% pooled | 9,920 | 48 |
| Value Line Cash | +2.1% | +2.9% | -2.2%mean -0.9% | 38.3% | 46.0%40% pooled | 6,204 | 27 |
| Owner Earnings | +3.9% | +2.2% | -2.0%mean +1.8% | 41.9% | 46.1%29% pooled | 5,315 | 44 |
| Buffettology Growth | +1.6% | +2.1% | -1.3%mean -0.4% | 42.4% | 52.3%29% pooled | 4,681 | 44 |
| Lynch Growth | -0.9% | +1.3% | -1.9%mean -2.2% | 41.3% | 45.6%23% pooled | 673 | 63 |
Methodology & caveatsShow
Ledger status
as of September 11, 2026 · cohorts May 16, 2026 → Aug 12, 2026 (71 verdict days) · live since May 16, 2026
How a number gets here
Each daily verdict anchors every company to its adjusted-close price on that date (no look-ahead). Once a horizon of 30, 90, 180, or 365 days has fully elapsed, we measure each company's total return to the next available bar and compare it to SPY over the exact same window. Companies without an anchor price or a matured end bar are excluded — never imputed. Deciles and quintiles are assigned over the full anchored cohort on the verdict day (not just the names that later survive to a matured bar), and companies with identical scores always share a bucket — assignment is deterministic and identical across horizons.
How the table is pooled
Across all matured cohorts we report the n-weighted mean return, SPY return, excess (return − SPY), and win-rate (share of names that beat SPY) for each PAS decile, pyramid tier, best-match guru, and PGS / PCI quintile. A bucket is shown only with at least 5 pooled observations — thinner samples are noise and are dropped. Tables also report cohorts (verdict days that contributed) beside n (names pooled).
Caveats — read before trusting
- Survivorship. Names that delist or stop trading lose their end bar and leave the cohort, so returns tilt slightly toward survivors. These are price returns of names that kept trading, not a tradable portfolio P&L.
- Short history. Snapshots began May 16, 2026; samples are small and noisy until cohorts accumulate. This is a forward record that grows in real time — not a backtest.
- New score families start at zero. PGS and PCI quintile cohorts accrue only from the day their snapshot columns shipped — earlier verdicts have no point-in-time record of those scores, so their history is never reconstructed after the fact.
- No look-ahead, but adjusted-close mixing. Base anchors and end bars share the adjusted-close basis, so there is no look-ahead. One residual: the base anchor is a true adjusted close while ordinary daily bars store the raw close as adjusted — at most a dividend-yield-scale (≈1–2%/yr) drift on individual names. The SPY benchmark is refreshed on a fully-adjusted basis, so the comparison line is unaffected.
- Not a recommendation. This measures historical price behavior of scored cohorts. Past performance does not predict future results, and nothing here is investment advice.
Why we grade the process, not just the outcome
Markets are a wicked learning environment: over short windows a sound process loses often and an unsound one wins often, so judging a method purely by its recent outcomes systematically rewards luck (Hogarth, Lejarraga & Soyer 2015, Current Directions in Psychological Science). That is why every claim on this page carries its sample size and a confidence interval, why verdicts refuse when the interval cannot carry them, and why a negative spread is published as plainly as a positive one.
The benchmark is SPY (S&P 500) on a dividend-adjusted basis — a deliberately unforgiving bar. Full methodology: docs/ACCOUNTABILITY_METHODOLOGY.md.
The full record lives in the app
Sign in for the interactive explorer — every horizon, per-bucket cohort histories, the PGS/PCI score-family tables, and each company's own since-first-verdict track record.