Sports research · test-before-claim · paper workflow
I test sports ideas with a written way to reject them.
At the August 21, 2026 ledger snapshot, zero ideas cleared the bar for real-money action. That is the system working, not failing.
Before a test runs, I write down what would change my mind. If the check fails, the idea stops.
Ledger snapshot: August 21, 2026, 11:30 UTC. These are recorded simulations, not live capital performance.
"This is what would prove me wrong." Written first, or it does not run.
The temptation in research is to notice a pattern and then search for the version of the test that lets the pattern survive. A falsifier is simply a written condition that ends the idea; it reverses that pressure.
Before a hypothesis enters the workflow, it gets a definition, a comparison, a time boundary, and a condition that would end the idea. The result is allowed to be boring.
The automation refreshes the data; I write the rule, inspect the result, and decide whether the idea advances.
Illustrative only, not a redacted real case: this is the shape of a falsifier before it enters the system.
Most hypotheses are rejected before they get anywhere close to a bet.
Of 814 logged questions, 3 have ever reached candidate status and 0 are active in this snapshot. 517 were rejected outright, most on process or data grounds, well before the question ever got a real look.
The first 517 break down into 393 process, 118 data, 4 clustering, and 2 suspicious decisions. The other 297 reached later states: 126 failed, 71 did not have enough evidence, 45 failed controls, 43 were replaced by a better version of the question, 9 had no clear answer, and 3 reached candidate status. Together those groups account for all 814 questions.
See the detailed rejection reasons
An idea can leave at any gate, not just the end.
Across 972 cumulative looks at the data, 43 hypotheses were replaced by a better version of the same question, 26 were replayed against fresh data, and 2 were formally disproved after initially looking promising. That last number is the one that matters most: a system that never rejects its own early candidates is not actually testing anything.
2,326,588 observations, 334 batches.
The research runs across 15 distinct mechanisms—the different ways a question is framed—and 6 broad market classes. 31 scheduled jobs keep the pipeline current; the capture is refreshed every 0.5 hours.
The ledger tracks paper bets, not real stakes.
28,007 paper positions are logged: 20,234 settled and 7,773 still open. Nothing here is sized for real capital. The ledger asks whether an idea survives new market data before it could ever justify risking money.
The interesting part is how often the answer was no.
The receipt behind the research result.
Open the research receipt
Counts and integrity metrics only. League and market names, candidate identities, edge and ROI values, thresholds, stake sizing, and closing-line values are deliberately excluded. The public page shows the checked summary; the raw working record remains private.
The claim
814 questions logged, with three ever reaching candidate status and none active in the snapshot.
The checks
Pre-registration, falsifiers, replay, controls, refutations, and the full rejection classification.
The scale
Observations, batches, paper-position status, and capture freshness.
Test what matters.
The same falsifier habit applies to any claim worth testing, sports, marketing, or otherwise.