Skip to content

231 · Overlapping 90-day books compounded monthly: the published returns were not portfolio returns (2026-09-30)

Status: SHIPPED as a restatement. Owner decision (2026-09-30): retract every CAGR, Sharpe, "10,000 EUR becomes X" figure and yearly returns table derived from the monthly top-10 harness, publish the portfolio-accounting figures below next to the STOXX Europe 600 on the same dates, and stop presenting the product as a market-beating strategy.

All figures below are MEASURED on 2026-09-30 (corpus read 12:14 UTC, strict serve pool 23,060 rows, composite snapshot configHash ed3c573192848099, code base origin/main e4fe5840) unless marked DEDUCED.

1. Finding

Two defects, both in the book runner shared by every published proof harness.

F1. Accounting. Each calendar month's book is the equal-weight mean of its top-10 picks' returnFromPub90d (a 90-day return) net of 0.6%. That series was then compounded 12 times a year and its Sharpe annualised with sqrt(12). A 90-day return was booked as if earned in one month, and the three books alive at any time were each given 100% of the capital. The CAGR came out about three times too high and the Sharpe inflated.

F2. Within-month look-ahead in selection. The top 10 were ranked over the whole calendar month, but each pick is entered at its own pubDate + 1. On the day a filing is published nobody can know whether it will still be in the month's top 10 at month end. Audit 200 listed this as L3. The trailing-percentile scorer removed the look-ahead from the score, not from the top-N cut.

Evidence (origin/main e4fe5840):

File What
scripts/_v13_bakeoff/pit-strict-2026-08-11.ts l.75-87, 513-537 sharpeAnn sqrt(12), cagr exponent 1/(m.length/12), iid monthly bootstrap; runTopN month bucket, bucket.slice(0, topN), mret += w*(winsor(r90)/100 - cost)
scripts/_v13_bakeoff/pit-scenarios-2026-08-11.ts verbatim copy, source of SCENARIO_RETURNS
scripts/_v13_bakeoff/pit-frame-implementation-2026-08-15.ts verbatim copy, source of SERVED_CONSTRUCTION_PROOF (cell F7)

Arithmetic check on the retired figures: the five SCENARIO_RETURNS.perYear multiples multiplied to 2.7071, the published finalMultiple 2.7069. The yearly table was the same 12x-compounded series cut by calendar year.

Secondary finding: the "top 10" is about five companies. On both cells 253 to 255 of the 510 picks repeat a company already in the same month's book (several filings by one issuer). There are 5.0 distinct issuers per monthly book on average (minimum 2). Diversification is about half of what "10 picks" implies, so every confidence interval below is too narrow.

2. Protocol

  • One runner. scripts/_v13_bakeoff/lib/nav-tranches.ts (simulation), lib/nav-stats.ts (NAV statistics, block bootstrap, the retired accounting kept for reconciliation), lib/entry-rules.ts (month-end top N, same-day threshold rule, distinct issuers), lib/overlap-cell.ts (one cell, every family). Unit tests colocated, including a toy case where the retired method reports more than four times the real CAGR. CLI: scripts/_v13_bakeoff/overlap-accounting-2026-09-30.ts, output overlap-accounting-2026-09-30.json.
  • Inputs. The two harnesses write their eligible pool (id, issuer, score, stored T+90 return, month-end top-10 flag) when POOL_DUMP is set, then stop. Nothing else in them changed. The runner re-derives the top 10 and checks it equals the harness's own (it does, on all three cells). T+30 and T+365 returns are read from BacktestResult inside a READ ONLY transaction. The index is EXSA.DE (iShares STOXX Europe 600) adjusted closes from Yahoo through fetchYahooDaily (src/lib/price-history.ts), 2,198 daily points to 2026-09-30.
  • Accounting. One pool of capital, marked to market every calendar day. Entry at pub + 1, exit at T+91. Each entry is sized at NAV/30, so a monthly book of 10 gets one third of the capital. Cost is split half at entry, half at exit. Endpoint returns are the harness's own winsorised returnFromPub90d, so only the accounting differs. Three selection variants:
    • same selection: the harness's 510 picks, all taken, gross exposure allowed above 100% (up to about 133%). Isolates F1.
    • capital-feasible: the same picks funded from cash only, a pick skipped when cash is exhausted.
    • no look-ahead (the headline): enter on pub + 1 if the filing's trailing-percentile score is at or above theta(M), the 120th-highest score among eligible filings published in the 365 days before month M starts (about 10 picks a month over the past year). Past data only. Capital-feasible, first come first served.
  • Path between anchors. PriceHistory reconciles too poorly with priceAtPub to use as-is (entry within 0.5% on 38% of picks), so each position's path goes through its T+30 and T+91 anchors, log-linear in between. CAGR depends only on the endpoints. Sharpe is an upper bound and MaxDD a lower bound in magnitude.
  • Statistics. Sharpe = (mean monthly NAV return - 0.25%) / sd x sqrt(12), the harness convention. CI95 by moving-block bootstrap (block 3, 2,000 draws, seed 42) on monthly NAV returns. CAGR over the real span, 2022-01-01 to the last exit (2026-06-27 or 2026-06-30).
  • Benchmarks, same accounting. (a) STOXX 600 on the same dates: every position replaced by EXSA.DE over its own entry and exit dates, no cost; computed for the headline entries and for the harness's 510 picks. (b) STOXX 600 bought once and held, 2022-01-01 to 2026-07-01. (c) Insider pool: every eligible filing bought, each monthly book a third of NAV, equal weight.

3. Results

3.1 Served construction, cell F7 (SERVED_CONSTRUCTION_PROOF, what production computes)

Window Retired (published) Same selection Capital-feasible No look-ahead STOXX 600, headline dates
ALL 51.2% / 1.58 / 58,027 EUR 14.0% / 1.34 [0.31, 2.54] / -11.4% 11.9% / 1.21 / -10.6%, n=457 10.2% / 0.99 [-0.16, 2.26] / -11.0%, n=441 9.9% / 0.65 [-0.01, 1.54] / -14.2%
W0 2022 48.0% / 1.37 11.0% / 0.67 9.6% / 0.61 5.4% / 0.28 0.6% / -0.12
W1 2023-24 22.3% / 1.10 6.5% / 0.75 6.5% / 0.79 5.7% / 0.54 10.6% / 0.96
W2 2025-26Q1 116.1% / 2.36 24.1% / 2.65 17.7% / 2.26 18.0% / 2.03 13.5% / 0.95
OOS30 2023-10 to 2026Q1 68.8% / 1.90 17.6% / 2.19 13.0% / 1.79 14.7% / 1.83 15.9% / 1.29

The retired harness on today's corpus gives 50.1% / 1.55 (published 51.2% / 1.58 on the 07:19 run of audit 230, 100 more pool rows). On the harness's own 510 picks the index returns 11.2% / 0.67.

3.2 Research cell (SCENARIO_RETURNS, PIT_STRICT_PROOF, formerly the home-page table)

Window Retired (published) Same selection Capital-feasible No look-ahead STOXX 600, headline dates
ALL 26.4% / 0.81 / 27,069 EUR 8.1% / 0.65 [-0.56, 1.85] / -11.2% 4.4% / 0.22 / -11.1%, n=460 6.6% / 0.59 [-0.57, 1.73] / -7.1%, n=382 8.7% / 0.64 [-0.05, 1.58] / -11.7%
W0 2022 35.2% / 0.95 9.5% / 0.57 6.7% / 0.38 1.4% / -0.27, n=51 -0.9% / -0.61
W1 2023-24 3.0% / 0.09 -0.3% / -0.60 -0.3% / -0.62 1.4% / -0.29 9.9% / 0.89
W2 2025-26Q1 66.1% / 1.38 17.9% / 2.04 9.0% / 1.10 17.6% / 2.19 12.2% / 0.83

The retired harness on today's corpus gives 26.1% / 0.84. On the harness's own 510 picks the index returns 11.1% / 0.66.

Calendar-year slices of the NAV (the retired yearly table is in the first row):

Year 2022 2023 2024 2025 2026 (to June)
Retired, published 35.2 16.2 -8.6 44.3 30.7 (Q1)
Research, same selection 3.7 8.3 -1.6 16.0 10.6
Research, no look-ahead 0.4 4.4 1.1 13.2 11.2
F7, same selection 5.9 11.9 6.3 26.8 12.9
F7, no look-ahead -0.2 13.5 3.5 20.1 10.0

3.3 Costs (no look-ahead, ALL, CAGR / Sharpe)

Cell 0.6% 1.5% 2.5%
F7 10.2 / 0.99 7.2 / 0.58 3.8 / 0.14
Research 6.6 / 0.59 4.1 / 0.19 1.3 / -0.25

At 1.5% round trip, closer to reality for Oslo and Stockholm small caps (DEDUCED, not measured per venue), F7 falls below the index on the same dates.

3.4 Benchmarks (ALL)

Benchmark CAGR Sharpe CI95 MaxDD
STOXX 600 buy and hold, 2022-01 to 2026-07 9.3% 0.51 [-0.14, 1.45] -20.7%
Insider pool, research cell, 0.6% -2.4% -0.79 [-2.05, 0.38] -16.2%
Insider pool, F7 cell, 0.6% -2.6% -0.79 [-2.11, 0.37] -17.4%

Reading. The ranking beats the pool of all eligible insider purchases by 9 to 13 points a year. Against the index on the same dates the served construction is level on return (10.2% against 9.9%) with a higher Sharpe (0.99 against 0.65), but its Sharpe CI95 starts below zero and the edge disappears at 1.5% cost. It trails the index in W1 2023-24 and leads it in 2022 and 2025-26. The research cell trails the index. No demonstrated edge over the market.

3.5 The 2026-08-16 flip, re-measured

B0 (pre-flip served construction) against F7 on today's corpus, no look-ahead: Sharpe 0.34 (DSR -0.08) against 0.99 (DSR 0.57). B0's stored pctOfMarketCap input is empty on this corpus (audit 230), so B0 no longer describes the pre-flip construction exactly. The retired-accounting delta (0.88 to 1.54, 2026-08-16) stays in flipDelta with its accounting named.

3.6 Reconciliation with the Python replica

A Python replica (scratch nav231.py, run the same morning on dumps from the same harnesses) is reproduced exactly on every point estimate: CAGR, Sharpe, MaxDD, trades and skipped counts, yearly slices, cost scenarios, pool and index figures. The eligible pools written today are identical to the replica's (same 15,167 and 13,807 ids, same scores, same returns, same top-10 flags). CI95 bounds differ by up to 0.09 because the replica drew its bootstrap with numpy's generator and the runner with a seeded mulberry32; the replica re-run with mulberry32 gives the runner's bounds exactly (F7 [-0.16, 2.26], research [-0.57, 1.73]).

The replica also estimated a market-factor path (Brownian bridge through the same anchors, market component from EXSA.DE, 10 seeds). That estimate was not ported to TypeScript: F7 no look-ahead Sharpe 0.64 central, MaxDD about -15.6% (Python only, not re-measured here). It is the more realistic Sharpe; the published 0.99 is the smoothed upper bound.

4. What changed

SSOT (values replaced, shapes kept, retired figures kept only in reconciliation fields):

  • src/lib/honest-backtest-proof.ts HONEST_BACKTEST_PROOF (new, re-exported as STRATEGY_PROOF.honestBacktest): the only figures the site may quote.
  • src/lib/metrics/scenario-returns.ts SCENARIO_RETURNS: perYear (NAV slices), capital, bySignalCount (top 5/10/20 = 5/10/20 a month under the threshold rule), byHorizon (per-pick means, today's corpus), meta.
  • src/lib/pit-strict-proof.ts PIT_STRICT_PROOF, src/lib/served-construction-proof.ts SERVED_CONSTRUCTION_PROOF: every window, plus stoxxSameDates, distinctCompaniesPerBook, insider pool; lookahead and flipDelta labelled retired-monthly-compounding.
  • src/lib/composite-proof.ts COMPOSITE_PROOF: status: "under-remeasurement", no longer displayed.

Public surfaces: every CAGR headline, "10,000 EUR becomes X" line and yearly returns table removed; one compact block shows the corrected backtest next to the STOXX 600 on the same dates, the one-line verdict and links to this audit and the live track record. Positioning moved to "the watch on European executive purchases, sorted and verified". The full list of files is in the pull request.

Not re-measured, same runner family, no longer displayed as returns: COMPOSITE_PROOF (audit 170), pitBacktest and oosResults (audit 169).

The ship gate in AGENTS.md ("Sharpe >= previous live version", DSR drop <= 0.30) has been judged on inflated Sharpes since at least audit 167. Decisions that compared cells inside one harness keep their ranking (same bias on both sides, DEDUCED); their levels were wrong.

5. Caveats

  1. Paths between anchors are interpolated: Sharpe values are upper bounds, MaxDD values lower bounds in magnitude (see 3.6).
  2. Survivorship: only priced rows enter; delisted issuers are missing, which biases returns upward (audit 200 §3).
  3. The no-look-ahead rule is one reasonable rule, not an optimised one. Its threshold needs the prior year of scored filings, so W0 2022 has few trades (research 51).
  4. Winsorisation at +/-50% is kept for comparability; it flatters fat left tails.
  5. The index is taken with zero cost and dividends reinvested (adjusted close); the picks are price returns. This favours the index by roughly the dividend yield times exposure (DEDUCED).
  6. About five issuers per monthly book: effective sample about half the stated N; intervals too narrow.

6. Next steps

  1. Re-run COMPOSITE_PROOF (audit 170 harness) through the same runner before it is shown again.
  2. Replace the anchor interpolation with adjusted daily prices (audit 190's Yahoo route) and add per-issuer dedupe.
  3. Re-judge the ship-gate history (audits 167 to 230) on corrected levels.
  4. Keep the live track record (/performance#suivi-reel) as the primary evidence; it needs no backtest accounting.

Reproduce

export DATABASE_URL="$(grep -m1 '^DATABASE_URL=' .env.local | sed 's/^DATABASE_URL=//; s/^"//; s/"$//')"
export DATABASE_URL="${DATABASE_URL/&pgbouncer=true/}"
mkdir -p /tmp/pools
POOL_DUMP=/tmp/pools/pool_research.json npx tsx scripts/_v13_bakeoff/pit-scenarios-2026-08-11.ts
POOL_DUMP=/tmp/pools/pool_frame npx tsx scripts/_v13_bakeoff/pit-frame-implementation-2026-08-15.ts
npx tsx scripts/_v13_bakeoff/overlap-accounting-2026-09-30.ts /tmp/pools
← All auditsSource: docs/method-review/231-overlap-accounting-2026-09-30.md