Sports research · test-before-claim · paper workflow

I test sports ideas with a written way to reject them.

At the August 21, 2026 ledger snapshot, zero ideas cleared the bar for real-money action. That is the system working, not failing.

Before a test runs, I write down what would change my mind. If the check fails, the idea stops.

The record 2 promising ideas failed my own follow-up checks. All 814 had a written rejection rule before testing. The two that looked promising still had to survive follow-up checks; that is the discipline, not the zero above it.
2,326,588 observations tested against
28,007 paper positions logged
100% written down first, with a way to be wrong
0 ideas cleared the bar for real-money action

Ledger snapshot: August 21, 2026, 11:30 UTC. These are recorded simulations, not live capital performance.

The starting point

"This is what would prove me wrong." Written first, or it does not run.

The temptation in research is to notice a pattern and then search for the version of the test that lets the pattern survive. A falsifier is simply a written condition that ends the idea; it reverses that pressure.

Before a hypothesis enters the workflow, it gets a definition, a comparison, a time boundary, and a condition that would end the idea. The result is allowed to be boring.

The automation refreshes the data; I write the rule, inspect the result, and decide whether the idea advances.

Every hypothesis in the receipt carries a commitment. The point is not to make the research look more scientific. It is to make the conclusion harder to move after the fact.

Illustrative only, not a redacted real case: this is the shape of a falsifier before it enters the system.

01 The idea. After a team plays a long midweek match, its next weekend game produces more goals than the usual baseline.
02 What would prove me wrong. If the pattern disappears after accounting for opponent strength and home advantage, or fails in a season kept out of the first test, the idea is dead.
03 The time limit. 500 observations, or the end of the season, whichever comes first.
04 What happened. If it fails the pre-written condition, it is rejected and does not get a second look. The rule exists before there is any reason to want the idea to survive.
Where ideas actually die

Most hypotheses are rejected before they get anywhere close to a bet.

Of 814 logged questions, 3 have ever reached candidate status and 0 are active in this snapshot. 517 were rejected outright, most on process or data grounds, well before the question ever got a real look.

The first 517 break down into 393 process, 118 data, 4 clustering, and 2 suspicious decisions. The other 297 reached later states: 126 failed, 71 did not have enough evidence, 45 failed controls, 43 were replaced by a better version of the question, 9 had no clear answer, and 3 reached candidate status. Together those groups account for all 814 questions.

01 · Written down814questions entered the research ledger
02 · Ever reached candidate3survived enough checks to earn that temporary label
03 · Active snapshot0remained active on August 21, 2026
517 stopped outrightAll three former candidates were later stood down; two failed my own follow-up checks after first looking promising.
The steep drop is the point. A written rejection rule should end most questions early, and a candidate must remain open to being disproved later.
See the detailed rejection reasons
Of 814 sports-research questions, most stopped for process, failed tests, or data quality; only 3 reached candidate status.
The detailed ledger separates three broad outcomes: stopped by process or data quality, stopped by a later evidence check, or still worth deeper review. The raw receipt keeps every internal status.
The integrity loop

An idea can leave at any gate, not just the end.

Across 972 cumulative looks at the data, 43 hypotheses were replaced by a better version of the same question, 26 were replayed against fresh data, and 2 were formally disproved after initially looking promising. That last number is the one that matters most: a system that never rejects its own early candidates is not actually testing anything.

Started with a written condition that would end the idea814 questions
Had timing scrambled to see whether the result could appear by accident695 checks
Reached the more expensive controlled-validation stage33 checks
Were replayed against fresh data26 questions
These rows show check coverage, not one funnel. Most questions stop before controlled validation because they already failed a cheaper process or data check. The deeper technical method injects a known test signal and compares it with scrambled results.
The scale underneath

2,326,588 observations, 334 batches.

The research runs across 15 distinct mechanisms—the different ways a question is framed—and 6 broad market classes. 31 scheduled jobs keep the pipeline current; the capture is refreshed every 0.5 hours.

Paper positions

The ledger tracks paper bets, not real stakes.

28,007 paper positions are logged: 20,234 settled and 7,773 still open. Nothing here is sized for real capital. The ledger asks whether an idea survives new market data before it could ever justify risking money.

Paper-position snapshot · August 21, 2026 20,234 of 28,007
20,234settled paper positions
7,773still open at the snapshot
$0 real dollarsThe scale belongs to a simulated ledger. It does not turn paper results into live performance.
The open and settled counts reconcile to the 28,007 total in the dated receipt.
What it took

The interesting part is how often the answer was no.

Time and scale
814 questions moved through a ledger with written commitments, comparisons, and dated follow-up.
What broke
Early ideas failed process, data, control, or replay checks; two that looked promising were later refuted.
Next time
I would make the rejection reason visible at the first gate, so a reader never has to infer why a question stopped.
Why it belongs
The system makes “not enough evidence” a recorded result, not a silent absence of a winner.
Evidence receipt

The receipt behind the research result.

Open the research receipt

Counts and integrity metrics only. League and market names, candidate identities, edge and ROI values, thresholds, stake sizing, and closing-line values are deliberately excluded. The public page shows the checked summary; the raw working record remains private.

The claim

814 questions logged, with three ever reaching candidate status and none active in the snapshot.

The checks

Pre-registration, falsifiers, replay, controls, refutations, and the full rejection classification.

The scale

Observations, batches, paper-position status, and capture freshness.

814 questions logged
3 candidates ever
0 candidates active
517 rejected total
2,326,588 observations
334 batches
15 mechanisms
6 market classes
100% pre-registration rate
100% falsifier rate
695 permutation-tested
33 controlled
26 replayed
43 superseded
2 refuted
972 cumulative looks
28,007 paper positions
20,234 paper settled
7,773 paper open
31 cron jobs
0.5 freshness hours
2026-08-21 11:30 UTC status generated