Why we publish what failed
2026-09-07
Most things we test do not beat holding Bitcoin. That is the normal result, and a research record that hides it is not a record, it is marketing. So the failures go out with the same weight as everything else, with the predictions we made before the run set next to what actually happened.
Here is the most recent batch, in full.
Batch 25: cycle and calendar models
The question came from a chart doing the rounds: cyclical timing models say Bitcoin is about to fall off a cliff. On inspection the chart was a Fibonacci extension from three prior tops with three log-scaled overlays fitted to match. No rule, so nothing to score. But the question underneath it was fair: does the ledger contain any cycle- or calendar-based model at all? It did not. Ten lines were written and pre-registered in the night of 6–7 September 2026 , all deterministic, at most two parameters each, no hindsight:
- the halving calendar, exiting 12 and 18 months after each halving;
- Pi Cycle, the 111-day and 350-day moving-average cross, with re-entry on the 200-day reclaim;
- the two-year-moving-average multiplier at two settings;
- a power-law trend fitted only on pre-ETF closes, with exposure set by the residual band, at two settings;
- a causal seasonality rule that stays out in a calendar month whose mean return over all prior years is negative;
- a 200-day moving-average switch as the naive control;
- hold-BTC.
Daily closes from September 2014, signal on the previous close, traded at the next. Fees included. Hold-BTC is 1.000.
What happened
Five of the eight cycle lines never fired after the ETF. Pi Cycle, both two-year-average bands and both power-law bands were fully invested from 2024-01-11 to 2026-08-31 without a single signal. On the benchmark epoch they are hold-BTC with a dormant switch: no record, not good, not bad.
The holdout is one fact, not a result. Bitcoin fell about 28 % between September 2025 and August 2026. Any line that happened to be out during that year passes the holdout. The two halving lines pass on exactly one trade each, after ranking eighth and ninth on the fit window. That is the “predicted a bear because the calendar said so” case, which is why the holdout is read as pass or fail and never used to rank.
Pi Cycle has the only long-history story: 3.43 × hold-BTC since 2015 on four trades, out in December 2017 and April 2021, back in February 2018 and May 2021. Two events in eleven years. Against a ledger of 301 trials its deflated Sharpe is 0.24, and its raw Sharpe sits below the maximum expected under the null for that many trials. It is a story about two dates, not a rule with a sample.
The seasonality rule topped the fit window at 1.092 and lost everywhere else: 0.24 since 2015, a fifth of it paid in fees. The known-bad line came first on the ranking window. That is a useful reminder that a fit window on its own can rank noise.
The two-year-average multiplier is a 2013-2017 artefact. All of its trades sit in one episode around November 2017; it has not fired since, not in 2021, not after the ETF.
The power law, fitted honestly on pre-ETF data, says Bitcoin has been below its long-run trend for the entire post-ETF era. As an exposure ladder it is a fee engine: over two hundred trades from the linear band.
The scorecard
Written before the run:
- Right: the two-year multiplier stays flat post-ETF; the halving lines finish below 1.000 on the fit window and post-ETF; the power law lands within 0.05 of 1.000 on the fit window; nothing earns a shadow book.
- Wrong: the halving lines were expected to fail the holdout on sample size, and both passed, because the holdout is a bear year and they happened to be out. The two-year multiplier was expected to beat hold-BTC since 2015; it did not. The power law was expected to be the likeliest holdout pass; it never traded. The seasonality rule was expected to sit at or below 1.000 everywhere; it came first on the fit window.
Four misses out of about eight calls. Every miss points the same way: these models do less than we assumed.
Verdict
No shadow book. Nothing ranks above hold-BTC on the fit window with a holdout pass. The ten lines stay in the ledger with a full window set, so the next time a chart says the cycle is about to turn, the comparison is already on record.
One thing to say plainly about the proof. The batch file, its result table and this verdict are hashed into Bitcoin and the rows are on the ledger — but the batch file was stamped at 23:39 UTC, five minutes after the run finished at 23:34. The predictions were written in the file before the run, and the file says so; the Bitcoin proof, for this batch, does not precede the outcome. Batch 25 was the first batch written after the ledger existed, and the stamp came at the end of the session instead of before the run. From batch 26 on the file is stamped before it runs, and a row whose proof is later than its run will be marked as such on the ledger.
What preparing this page found
Writing the scoreboard exposed a structural flaw that a week of running had not. The nineteen shadow books were meant to be the live book with one change each; they turned out to be running on the engine’s default universe and their own paper base, and the invariant meant to keep the twin equal to the live book compared raw balances, so it went blind the moment the live capital changed on 2 September. Reading the code the next session found more: the books were built on an older default configuration of the brain, not the live one, and each was told it held what the live book held. All nineteen were restarted on 2026-09-07 from the live book’s holdings with the live parameters, the invariant now compares multiples, and its first reading after the restart was zero to six decimals. The pre-reset week is archived. The table is on the scoreboard with that stated above it. It is on record here because it is the kind of thing a private research process quietly fixes and a public one has to say.
Why this is worth publishing
A batch that produced nothing usable still produced something: the denominator. The 301st trial makes the 302nd harder to pass by chance, and only a public ledger of the failures lets a reader check that the denominator is real. If, one day, a line does pass everything and runs 90 days as a shadow book and gets promoted, the value of that claim rests entirely on how many times we said, in advance and in public, that something would not work — and were right, or admitted we were wrong.