Options research · paper only · evidence gates

I test options ideas before they reach real money.

Every idea is written down, tested, checked, and kept on paper until the evidence is strong enough.

I built a paper-only system to challenge promising ideas. It records each question, tests it against history and controls, and stays away from real money until the evidence is strong.

The record 446 verdicts recorded; 11 cleared the research gate; the system remains on paper. Not one dollar of real money has ever been placed. The system stops when the evidence is incomplete. Decision: STAY ON PAPER.
15,412 historical test results
11 passed the research gate
220/583 of the planned test combinations have been run
11 incidents logged and retained
The ordinary-language version

An idea has to survive the whole loop.

01Write the question. Say what the idea is supposed to explain.
02Decide what would prove it wrong. Set the comparison and failure rule before running the test.
03Run the historical check. Replay the question against past data and keep the underlying records.
04Challenge the result. Compare it with a simpler explanation and keep the trade on paper.
05Record what happened. Pass, fail, noise, or not enough data all stay in the ledger.
06Reconcile before the next run. Check what changed and keep real-money decisions out of the loop.
The problem

Finding something interesting is easy. Knowing whether it is real is the job.

A historical test can produce a beautiful line and still answer a question nobody wrote down in advance. The danger is not that the result is ugly. It is that the result becomes persuasive before anyone asks what produced it by accident.

This project makes the challenge part of the system: finite questions, explicit controls, paper execution, and a record of what got rejected. Markets are the harshest place to learn that a persuasive chart can be nothing. That is the same lesson a campaign dashboard teaches, only faster and with a scoreboard.

The finding

Most ideas do not survive contact with a control group.

01 · Recorded446research verdicts written into the ledger
02 · Still plausible38were not negative or incomplete after the first cut
03 · Passed11cleared the full research gate
408 stopped earlyAnother 27 looked better or repeated successfully but still did not qualify as a clean pass.
Passing is a high bar on purpose. A result can look interesting, beat the comparison, or repeat on a second run and still not clear the research gate.
446 · RECORDED 91.5% come back negative or incomplete. Of 446 verdicts logged, 408 failed, looked like noise, or did not have enough data to judge. Only 11 passed the research gate.
STATUS · 2026-08-16 20:33 The operating decision stays conservative on purpose. Decision at the receipt snapshot: STAY ON PAPER. Nothing here has been sized for real capital.
Open the detailed verdict distribution
Most of 446 options-research verdicts were noise, unusable data, or failed tests; 11 passed.
The full ledger keeps every outcome visible: better, passed, repeated, weaker, noise, failed, or no usable data.
The scale of the check

The verdicts rest on a much larger evidence set.

15,412 historical-test results and 988,538 individual simulated-trade records fed into those 446 verdicts, over 6,479.2 hours of computer time. The verdicts are the conclusions; these are the records underneath them.

Final research conclusions446 verdicts
Results from historical simulations15,412 results
Individual simulated trades inspected underneath the results988,538 records
6,479.2 computer hoursRun time is evidence of scale, not proof that an idea works. The gate still decides what the system is allowed to claim.
These are different layers of the same research process, not one funnel and not 988,538 separate experiments.
Declared research plan covered so far 220 of 583
220planned test combinations with evidence
583planned test combinations in the full map
Untested combinations are not evidence. Showing the full denominator makes the unfinished work visible.
System checks

The research is only as trustworthy as the system running it.

A conclusion is only as good as the steps that produced it, so those steps are checked on a normal cadence—not just when a result looks surprising.

TESTS 952 passing, 48.8% line coverage. Not full coverage. Line coverage on a research engine is a floor, not a target.
INCIDENTS 11 logged, median time-to-detect 24.0 hours. Each incident stays in the internal log with what broke and when the system caught it. This portfolio shows the count and one sanitized example because hidden failures do not make a research engine more trustworthy.
ONE INCIDENT, SANITIZED A broken run stayed out of the verdict ledger. A scheduled run produced incomplete evidence, so the health check stopped it before a conclusion could be recorded. After the missing coverage was repaired, the run was repeated and the incident remained in the log.
AUTOMATION 34 scheduled jobs keep the pipeline moving. The research can run on a normal cadence without the owner watching every stage constantly. That changes the workload, not the decision boundary.
What it took

The useful result is the refusal to overread a pass.

Time and scale
15,412 historical-test results, 988,538 individual simulated-trade records, and a recurring operating cadence made the evidence expensive enough to respect.
What broke
One incomplete run was stopped before it could enter the verdict ledger; the incident stayed visible after repair.
Next time
I would surface coverage gaps even earlier and keep the “not enough evidence” path as prominent as the pass path.
Why it belongs
The system makes a conservative decision—stay on paper—while showing exactly how the checks and failures shaped it.
Evidence receipt

The receipt behind the research result.

Open the research receipt

Counts and process metrics only. Strategy names, underlyings, thresholds, position sizing, and every positive edge value are deliberately excluded. The public page shows the checked summary; the raw working record remains private.

The claim

Verdicts recorded, the research gate, and the paper-only governance decision.

The checks

Controls, incidents, coverage, and the sanitized failure that stayed out of the ledger.

The scale

Backtest rows, matrix coverage, engine time, and operating cadence.

446 verdicts recorded
11 passed the gate
408 negative total
91.5% negative share
15,412 backtest results
988,538 per-trade rows
6,479.2 engine hours
220 matrix cells tested
583 matrix cells planned
952 tests passing
48.8 % line coverage
11 incidents logged
24.0 median MTTD hours
34 cron jobs
STAY ON PAPER governance verdict
2026-08-16 20:33 status generated
If this discipline is useful to you

Let evidence decide.

The system is built to say "not enough evidence" without flinching, a habit worth having on any team.