Marketing measurement · plan written in advance · public dataset
Late-arriving conversions made reported cost per customer look about 4× too low.
A public-benchmark study of conversions credited after the budget decision.
I used a public Criteo advertising dataset to test whether late-arriving conversions could make an old report look better than the evidence available at the time. I wrote the test plan before looking at the result.
The 60.3% is a follow-up description, not part of the original plan: it is the share of test budget behind campaigns that failed the timing check.
A conversion recorded today may belong to last week's budget decision.
Nothing necessarily broke in tracking. A report can receive a conversion after the budget decision and file that credit into the older campaign story. That timing gap is where the problem hides: the campaign looks stronger than the evidence available when the decision was made.
This study measures that gap directly on a dataset anyone can download.
The human work is choosing the cutoff, writing the rule down before seeing the result, and deciding what the evidence allows next.
Score the old report, rebuild the decision-time view, then choose what deserves another test.
The leak is real, it is large, and it is not random noise.
A correction matters only if it changes the next budget decision.
The real test of a correction is whether it changes what you would have done. On public-benchmark data held back from the analysis, the first rule selects 274 campaigns at a cost per customer of 0.01782. The corrected rule selects a different, smaller set: 99 campaigns at 0.00877, a 50.8% improvement in this benchmark comparison, not proof of live campaign lift.
One anonymous campaign reports a cost per customer of 0.00828 before the correction and 0.05076 after: a 6.1× difference caused by 431 late-arriving credits. These are the benchmark's own units, not dollars. The source identifier stays in the receipt for auditability.
Two implementation bugs were caught before the result was treated as evidence. One check was first compared against the wrong baseline, and a second miscounted which campaigns belonged in the test group. Both were fixed, everything downstream was re-run, and the original plan stayed intact.
Does this happen in your account?
Probably, and it is not anyone's fault. Many ad platforms backfill conversions into the day of the click, not the day the conversion occurred. That is the right choice for understanding history and the wrong one for judging a decision you already made.
Read the evidence as a set, not a hero chart.
The finding above rests on more than one chart. This is the placebo run referenced in The finding: the test that had to come back empty for the result to mean anything.
The complete benchmark ledger, for readers who want the audit trail.
Open the measurement receipt
Every value here is generated from the public benchmark receipt. Nothing is redacted, because the underlying data is already public. The public page shows the checked summary; the raw working record remains private.
The attribution-credit count is grouped by campaign-day cohort; it is not a count of unique customers or revenue.
The claim
The first figures describe what changed: candidates before and after the time gate, the decision ledger, and the registered out-of-sample comparison.
The checks
The next figures test leakage, attribution-model sensitivity, placebo behavior, and weekly stability.
The scale
The final figures describe volume, cohort coverage, and the worked example. Units stay with each value below.
The result only became useful after two uncomfortable checks.
The full method and evidence trail live on a separate page.
This page stays readable as a summary. The method page carries the plan written before analysis and five supporting charts, the credit-assignment and stability checks, and a clear boundary around this public-benchmark study. The source write-up remains separate from this portfolio page, and publication remains gated until the final source-release, redaction, and licensing checks pass.
Find the leak.
This is the closest match to my day job. I’d rather find the leak in your data than in someone else’s.