Marketing measurement · method & checks · optional

How I tested whether campaign reports counted late-arriving conversions.

Four checks test the timing rule, the way credit is assigned, the risk of testing many campaigns at once, and week-by-week stability.

This page asks a narrower question: does the finding survive a different timing rule, a different way of assigning credit, a many-comparisons check, and separate weekly tests? The plan written before analysis, supporting charts, and source record live here.

The finding The timing result survived a second check with 999 additional scrambles. The first check scrambled the timing 20 times under a plan written in advance. The confirmation repeated that challenge 999 more times.
Before the detail

Four checks, translated.

Attribution modelThe rule that decides which campaign gets credit for a conversion.
False-discovery-rate correctionA guard against mistaking random winners for real findings when many campaigns are tested at once.
Permutation testA comparison made by scrambling the timing so we can see what a no-effect result would look like.
Walk-forward checkA week-by-week test that uses the information available at each point instead of looking back with the full dataset.
Does the model choice matter?

Last-click and multi-touch change the answer for some campaigns.

Two attribution models were run on the same corrected data: 100 rows scored last-click, 81 scored multi-touch. 45 campaigns flip verdict depending on the model, a real methodological choice, not a rounding difference.

Testing many campaigns at once

The correction used 616 eligible campaigns, not all 675.

The benchmark contains 675 campaigns, but 59 fell below the held-out impression floor, leaving 616 eligible for this test family. The uncorrected screen had 530 apparent winners. At q = 0.1, Benjamini–Hochberg requires p < 0.002; 480 of the 616 clear that bar. That survival rate reflects the volume in each eligible campaign, not a loose correction.

Does it travel?

A walk-forward check across weekly windows.

This is a separate check from the leakage screen above. Of the 652 campaigns with enough weekly volume, 93 stayed consistently above baseline, 473 stayed consistently at or below it, and 86 flipped sides. Read those first two numbers as one result, not opposing ones: 87% had a stable weekly signal, whichever side of baseline they landed on. Week-one agreement with the full-sample verdict was 88.4%, which is the honest result for a real dataset, not a perfect one.

The prediction written before analysis did not fully match what happened. The first eligible week agreed with the full-sample verdict 88.4% of the time. I kept that original weekly prediction in the record instead of rewriting it after seeing the result; the mismatch is part of the evidence, not an inconvenience to hide.

Who is left after the correction?

The campaigns that survive skew toward a specific profile.

Surviving campaigns run at 2.62× the impression volume, 1.27× the click-through rate, and 0.16× the cost per customer of the ones that do not survive. Their volume sits around the 77.3% percentile—higher than roughly three quarters of campaigns. The correction is not one-size-fits-all; it has a shape, and that shape is worth knowing before applying it elsewhere.

Supporting charts

Five checks, read as a set rather than in isolation.

Each asks a different question: does the finding hold over time, weaken with thin data, depend on one attribution model, survive many comparisons, and match the timing mechanism?

Of 652 weekly campaign checks, 473 moved in the expected direction, 93 moved the other way, and 86 were mixed.
Does it hold over time? Each week is scored separately rather than pooled. Most weekly checks move in the expected direction; mixed and opposing weeks remain visible instead of being averaged away.
Uncertainty narrows across five campaign-volume groups, from 0.746 percent in the smallest group to 0.141 percent in the largest.
Where is the estimate thin? Low-volume campaigns have wider uncertainty: 0.00746 in the smallest-volume group versus 0.00204 in the comparison group used by the original check. The chart shows all five volume groups, not only two endpoints.
Changing the attribution model changed the verdict for 45 campaigns; 100 were retained under last-click and 81 under multi-touch.
Does one attribution model decide the answer? Switching from last-click to multi-touch changes the verdict on 45 campaigns. That sensitivity belongs in the decision, not in a footnote.
A false-discovery correction reduced the apparent winner count from 530 to 480 across 616 eligible campaigns.
How many apparent winners survive a many-comparisons check? The deeper method uses 616 eligible campaigns. At q = 0.1, the p < 0.002 threshold leaves 480 of the 530 uncorrected winners.
Only 39 percent of credited conversions were visible at the budget decision; 61 percent arrived later.
Does the proposed mechanism exist? Only 39% of credited conversions were visible at the decision; 61% arrived later. That is the room in which the reporting problem can happen.
Provenance

A real, tested pipeline, not a one-off script.

The study runs as a 21-script pipeline against a public Criteo advertising dataset: seven ordered stages from capture through decision, then separate scripts for the placebo, weekly stability, and other documented checks, plus a test suite that must pass before a number is trusted. The source dataset, decision-run identifier, and schema-verification status are all in the receipt below. This is a public-benchmark study, so nothing here is redacted. The public page shows the checked summary; the raw working record remains private. The analysis plan was written before the placebo results were known. During the final clean rerun, a provenance-order concern was corrected and every later output was regenerated without rewriting that plan.

If you want this level of rigor

When claims matter, go deeper.

A plan written before analysis and a public benchmark are what it takes to trust a result. That standard travels to any measurement problem.