Beyond the Dashboard · Nº 10

Beyond the Dashboard

A field guide to experimentation & statistical judgment

Every experiment is an attempted refutation of the claim “this change does nothing,” and the statistical machinery — power, p-values, intervals, corrections — exists for one purpose: to keep you from fooling yourself. The ways teams fool themselves are finite, nameable, and mostly self-inflicted. Your job is not to run the math. It is to be unfoolable by it.

Module 01 Why correlation keeps winning

Every causal claim your organization makes is a comparison against a world that does not exist: the same customers, the same quarter, without the change. You never observe that world. Everything in this guide — randomization, power, intervals, difference-in-differences — is a strategy for constructing a credible stand-in for it, and every failure of judgment in the field is a bad stand-in accepted without argument.

This module is about the three ways a bad stand-in gets accepted. Each has a name, a mechanism, and a tell you can spot in a slide. You are already trained for this work: a correlational claim is circumstantial evidence, and your instinct to ask who selected the exhibits, what else could explain them, and who is missing from the room is exactly the right instinct. What follows gives that instinct a vocabulary precise enough to argue with an analyst.

The counterfactual question

Meridian, a subscription-billing company, ships Draft Assist — an LLM feature that pre-drafts replies for its 240 support agents. In March, mean resolution time falls from 262 minutes to 241. The deck says the feature saved 21 minutes per contact. It says no such thing yet, because the claim being made is about a comparison, and only half of the comparison has been observed.

The counterfactual is the outcome the same units would have had, over the same period, without the treatment. It is unobservable by construction — you cannot both ship and not ship to the same contacts in the same week. This is the fundamental problem of causal inference, and it has exactly one class of solution: find or build a group whose observed outcome is a credible substitute for the missing branch.

A single population forking into an observed shipped world and an unobserved counterfactual world, with the causal effect drawn as the gap between the two endpointsMeridian contactsweek 0, before the forkWorld A — Draft Assist shipped (observed)World B — not shipped (never observed)Observed241 minCounterfactual? minthe causaleffect livesin this gap
Figure 1.1 — The gap you can never measure directly. A causal effect is the difference between an outcome you observed and an outcome that never happened. Since the dashed branch is permanently unobservable, every method in this guide is a way of manufacturing a defensible stand-in for it — and every causal dispute is a dispute about that stand-in.

Notice what the March dashboard actually substituted for the dashed branch: February. That substitution assumes nothing else changed between the two months — no seasonal dip in dispute volume, no staffing change, no queue-routing update, no post-holiday billing surge working its way out of the system. Before/after comparison is not a weak experiment. It is a strong assumption wearing a chart.

The load-bearing idea

“Compared to what?” is the first cross-examination question, and it has only three honest answers: a randomized control group, a designed quasi-experimental comparison whose assumptions you are prepared to argue (module 7), or “we do not know.” Last month is not one of them.

Confounding

A confounder is a variable that causes both the supposed cause and the supposed effect. It manufactures correlation out of nothing causal, and it does so with total statistical realism: the correlation is genuinely there, reproducible, and significant at any sample size you like.

Meridian’s growth team observes that customers who adopt the mobile app retain at roughly twice the rate of those who do not, and proposes a campaign to drive app installs. Work the mechanism. Engaged customers — the ones who log in, configure things, care about the product — both install the app and stay. Engagement causes adoption; engagement causes retention. The observed adoption–retention correlation flows entirely through those two edges. Drive installs among the disengaged and you get installs, not retention.

Causal diagram showing engagement as a confounder of app adoption and retention, with the backdoor path highlightedthe confounderEngagementApp adoptionRetentioncausal edge — may be ≈ 0backdoor path: adoption ← engagement → retentionthe 2× correlation can travel entirely along the dashed route
Figure 1.2 — The backdoor path. Correlation between adoption and retention can travel two routes: forward along the causal edge, or backward out of adoption, up through the confounder, and down into retention. That second route is the backdoor path. Blocking it — by randomizing adoption, or by conditioning on engagement if you can measure it fully — is what converts an association into a causal estimate.

The phrase to distrust is “we controlled for that.” Statistical adjustment closes a backdoor path only for confounders you measured, measured well, and thought of in advance. Engagement is a construct; your proxy for it is logins per week, which captures perhaps half of it. The unmeasured half keeps the backdoor open, and no amount of regression output tells you how wide.

From your other life

A confounder is the alternative explanation the opposing party gets to argue. Your adjustment is your rebuttal to the explanations you anticipated. Randomization is different in kind: it forecloses every alternative explanation at once, including the ones nobody thought to raise — which is why it is a guarantee and adjustment is only an argument.

Selection bias

Confounding is about a third variable. Selection bias is about the door: who ends up in your dataset is itself an outcome, and if the treatment influenced who walked through, the comparison is no longer between comparable groups.

Meridian ran Draft Assist as an opt-in beta first. The 38 agents who volunteered showed CSAT six points above the rest of the floor. The obvious read is that the feature raises satisfaction. The other read is that agents who volunteer for a new AI tool skew senior, skew engaged, and skewed high on CSAT before the beta existed. Seniority causes both opt-in and CSAT; the beta–CSAT association is manufactured by who chose to be measured.

Causal diagram of selection bias: seniority drives both beta opt-in and CSAT, and the analysis is conditioned on opt-in, producing a non-causal associationcommon cause of bothAgent seniorityBeta opt-inCSATobserved +6pp gap — no arrow, no causethe analysis lives inside this box:volunteers onlyComparing volunteers to non-volunteers measures who volunteered.
Figure 1.3 — Selection bias as a filter on the sample. No arrow runs from beta opt-in to CSAT, yet the two are associated in the data, because a common cause drives both and the analysis is conditioned on the opt-in door. The gold box is the tell: whenever the analyzed population was assembled by a process the treatment could influence, the comparison is between kinds of people, not between conditions.

A subtler version conditions on something downstream of the treatment. “Among escalated tickets, drafted replies resolve slower” sounds like an indictment of Draft Assist. But escalation is itself affected by the feature: if drafts resolve the easy cases cleanly, the drafted tickets that still escalate are the genuinely hard residue, while the undrafted escalations include cases that escalated for want of a good first reply. You filtered on an effect of the treatment and compared what was left.

The tell

Any comparison where membership in the analyzed group was chosen — by users, by agents, by a filter applied after the treatment — is a selection comparison until proven otherwise. Ask: what process decided who is in this dataset, and could the treatment have touched that process?

From your other life

This is the sampling objection: the exhibits were selected by the party offering them. You would never accept “here are the twelve emails we chose to produce” as a fair picture of the correspondence. An opt-in beta is the same document production, run by volunteers.

Survivorship

Survivorship bias is selection with a timer. The analysis runs on units that lasted long enough to be measured, and the ones that failed are not merely underrepresented — they are structurally absent, invisible to a query that starts from the current customer table.

Meridian’s customer-success lead observes that the longest-tenured accounts nearly all onboarded with white-glove setup, and proposes white-glove for everyone. The churned white-glove accounts are not in the room. If white-glove was disproportionately given to large, complex accounts, and complex accounts churn harder when they churn at all, the surviving population can look like a white-glove endorsement while the full cohort shows nothing.

A cohort funnel narrowing over 24 months, with churned accounts exiting downward and the analyst observing only the survivors at the right edge1,200startWhite-glove onboarding cohort, months 0 → 24790 churned — absent from every query that starts at “current customers”what the analyst sees410 accounts“all white-glove!”
Figure 1.4 — The camera points only at survivors. A cohort of 1,200 narrows to 410 over two years; the 790 exits leave the frame entirely. Any statement of the form “our best customers all did X” is computed inside the gold box, where the failures of X have already been deleted, so the statement carries no information about whether X helps.
The canonical form

Wartime analysts examined returning bombers, mapped the bullet holes, and proposed armoring the areas most hit. Abraham Wald pointed out that the sample consisted entirely of planes that made it back: the unhit regions were where the missing planes had been hit. Armor the engines. The tell is identical to the white-glove claim — the data was collected from survivors, so the pattern of damage is a map of survivable damage.

The operational tell is a query shape: any claim computed from current customers, current employees, retained cohorts, or shipped features carries survivorship risk. The corrective is to define the denominator before you look at the numerator — the full cohort as of its entry date, exits included — and to notice when you cannot, because the exits were never logged.

Module 02 The anatomy of a valid experiment

Module 1 ended with a demand: to license a causal verb you need a credible stand-in for the counterfactual. Randomization is the only method that builds one by construction rather than by argument. This module is about what randomization actually buys, the four decisions that determine whether you collect on it, and the two failure modes — unit mismatch and interference — that void the purchase without producing any visible error.

The through-line is the Draft Assist experiment at Meridian: 240 agents across billing, cancellations, and disputes queues, roughly 12,000 contacts per week. By the end of this module you will have the design one-pager for it, minus the sample size, which module 3 computes with real arithmetic.

Randomization as the confounder-killer

Randomization assigns units to arms by a mechanism unrelated to anything about the unit. That single property does the work: because the coin does not know an agent’s tenure, an account’s size, or a contact’s difficulty, no such variable can systematically differ between arms except by chance you can quantify. Every backdoor path from module 1 is severed at once — the measured ones, the unmeasured ones, and the ones nobody has thought of.

Stated formally this is exchangeability: the arms are interchangeable in expectation before treatment, so an outcome difference afterward has exactly one systematic explanation left. Stated practically: randomization is the only technique in this guide whose validity does not depend on your imagination.

The load-bearing idea

Adjustment handles the confounders you named. Randomization handles the confounders you did not name — which, since you cannot enumerate what you failed to think of, is the entire reason it sits at the top of the evidence hierarchy.

Randomization is a mechanism, and mechanisms break. The calibration ritual is the A/A test: run the full pipeline with both arms receiving the identical experience, and confirm the analysis finds nothing at roughly the advertised rate. An A/A that produces significant differences on 3 of 10 metrics is telling you the bucketing is correlated with something — assignment sticky by session, a caching layer serving one arm stale content, a logging join that drops rows unevenly. Better to learn that on a null test than to spend it on a real one.

The unit decision

The unit of randomization is the entity the coin flip lands on. For Draft Assist there are three candidates, and the choice among them determines the experiment’s power, the integrity of its exposure, and its vulnerability to contamination — all three at once, in tension.

Contact-level. Each incoming contact independently gets drafts or not. Maximum statistical resolution: 12,000 independent-ish observations a week. But an agent works both kinds of contact in the same shift, so treatment is not cleanly isolated to the treated units — an agent who learns better phrasing from drafts carries that phrasing into their undrafted replies.

Agent-level. Each agent is drafts-on or drafts-off for the whole run. Exposure is clean and the treatment matches how the feature would actually ship. But outcomes within one agent are correlated — a fast agent is fast on all 300 of their contacts — so 240 agents are worth vastly less than 24,000 contacts, an arithmetic module 3 puts a price on.

Queue-level. Whole queues are assigned. Contamination is nearly eliminated, and so is any hope of precision: three queues means an effective sample of three.

Grid comparing contact, agent, and queue randomization across statistical power, exposure integrity, and interference risk, with the chosen design markedUnit of randomizationStatistical powerExposure integrityInterference riskContactchosen designhigh — 12,000/weekpartial — agent seesboth conditionslearning spillover,biases toward nullAgentlow — effective n capsnear 2,000/armcleanshared queue stillcouples the armsQueueeffective n = 3cleanlowThe voided cell: randomize by agent, analyze by contact.Standard errors computed as if 24,000 independent observations — significance fabricated.
Figure 2.1 — The unit decision, priced across three axes. No unit wins on all three: contact-level buys power and pays in exposure purity, agent-level buys clean exposure and pays in effective sample size, queue-level buys isolation and pays everything. The dashed band is not a fourth option but the error that voids any of them — analyzing at a finer grain than you randomized.

The binding resolution. Draft Assist randomizes at the contact level. Drafts appear only on treated contacts; the agent-learning spillover is documented as a threat, and its direction reasoned out rather than waved at: an agent who improves from exposure to drafts improves on their untreated contacts too, which lifts the control arm and shrinks the measured gap. The bias runs toward the null, so a win survives it and a null result is genuinely ambiguous. That asymmetry is what makes the choice defensible. Module 3 shows the alternative arithmetic: the agent-level design cannot answer the 5% question at any run length, which turns a preference into a constraint.

Anti-pattern — the unit mismatch

Symptom: the design doc says “randomized by agent” and the readout says “n = 24,000 contacts, p = 0.01.” Mechanism: contacts within an agent are correlated, so they carry less information than independent observations; treating them as independent shrinks the standard errors. What it fabricates: significance. Not a biased estimate — a confidence interval several times too narrow, and a p-value with no relationship to any error rate. Corrective: analyze at the unit of randomization (agent means), or use standard errors clustered by agent, and say which in the readout.

Interference: when the arms touch

Every experiment carries an assumption so quiet it usually goes unstated: my treatment does not affect your outcome. Violations are called interference (SUTVA violation), and they are the most common way a technically flawless experiment measures the wrong thing.

At Meridian the mechanism is the shared queue. Treated agents resolve contacts faster, so they return to the queue sooner and pull more work. The queue drains differently. Control agents now face a shorter backlog, a different mix of aged versus fresh contacts, and less time pressure. Their resolution time changes — not because of the feature, but because of other people’s feature. The control arm has stopped estimating the no-feature world.

Swimlane diagram showing treated agents and control agents coupled through a shared work queue, with the intended causal path solid and the leak path dashedTreatment armShared stateControl arm120 agentsdrafts onresolve fasterone work queue · 12,000 contacts/week · shared backlog and aging120 agentsdrafts offoutcome shifts anywayno longer the no-feature worlddrains the queue fasterchanges what control agents face
Figure 2.2 — The leak path. The intended path (solid) runs from treatment to treated outcomes. The leak (dashed) runs through shared state: faster treated agents drain the common queue, which changes the backlog, aging, and pace the control arm experiences. Once the arms are coupled, the treatment–control difference estimates neither the feature’s effect nor zero — it estimates the effect of the mixture.

Reason out the direction rather than asserting it. If control agents face a lighter, fresher backlog, their resolution times improve, the gap narrows, and the experiment understates the feature. If instead the treated agents skim the easy contacts and leave control agents a harder residue, control times worsen and the experiment overstates. Which mechanism dominates is an empirical question you can partly answer — compare control-arm contact difficulty and queue age against the pre-period.

The mitigations, in ascending cost: measure the spillover (track control-arm queue characteristics as a diagnostic), isolate routing so arms draw from separate contact pools, or randomize at the queue or site level and accept the power loss. Meridian takes the first: contact-level randomization means the queue is shared by construction, so the design documents the coupling, monitors the control arm’s workload composition, and treats the null-ward bias as a conservative feature of the win case.

From your other life

Interference is a side channel. Two processes are supposed to be isolated; they share a resource; state leaks through the resource rather than through the interface anyone is watching. You already know how the analysis goes — enumerate what the two arms share, and treat each shared resource as a candidate channel until you have argued it is not one.

Exposure, eligibility, and the analysis population

Assignment is not exposure. Of the contacts assigned to Draft Assist, some fraction of agents will never open the drafted reply — busy, skeptical, or working a contact type where the draft is obviously useless. Those contacts were treated by the coin and untreated in fact, and what you do about that determines what question your experiment answers.

Intention-to-treat analyzes every unit by its assignment, regardless of what it actually experienced. This feels wrong to engineers and right to lawyers: it deliberately includes the failures of adoption in the estimate. It is also the only analysis that preserves what randomization bought, and it answers the question you actually face — what happens to resolution time if we launch this feature to the floor, including the agents who ignore it?

The seductive alternative is per-protocol: analyze only the contacts where the draft was opened and used. That comparison is not randomized. Agents choose which drafts to use, and they choose the good ones on the tractable contacts. You have reconstructed, precisely, the self-selection that module 1 spent thirty minutes killing — with the added indignity that the number will look better, which is why someone will ask for it.

What dilution does to your arithmetic

If only 70% of assigned contacts are genuinely exposed, the intention-to-treat estimate is roughly 70% of the effect on the exposed. Plan for it: an effect worth 5% among users shows up as 3.5% overall, and an MDE set at 5% will miss it. Either set the MDE against the diluted effect or fix adoption before testing. Discovering dilution after the run is discovering you designed for the wrong question.

Eligibility is the third piece and belongs in writing before launch: which contacts can enter the experiment at all. Meridian excludes contacts routed to the fraud-review queue (different tooling, different SLAs) and contacts opened before the launch timestamp. The rule matters less than its timing — an eligibility filter written after the data exists is a subgroup selection, which module 5 names and prices.

The design one-pager

Every discipline in the rest of this guide depends on one artifact existing before the data does. Without it, module 5’s corrections are unenforceable, because there is no record of what was promised. The design one-pager is short by intention — six commitments, one page, circulated and dated.

Pipeline from population through eligibility, randomization, exposure, and outcome window to analysis, with the six design commitments pinned to the stages they governPopulationall contactsEligibilityfilterRandomizeunit: contactExposuredraft shownOutcomewindowAnalysis1 · eligibility rule2 · unit + split3 · exposure def.4 · run length5 · primary+ MDE6 · guardrails, with veto authority — and the sample-ratio-mismatch check before any metric is readCSAT (78%) · reopen rate (4.1%) · compliance-flag rate (0.8%)If the observed split departs materially from 50/50, stop: the assignment machinery is broken and nothing downstream is readable.
Figure 2.3 — Six commitments pinned to the stages they govern. Each commitment binds one stage of the pipeline before data exists: eligibility governs entry, unit and split govern assignment, the exposure definition governs what counts as treated, run length governs when you look, primary metric and MDE govern the verdict, and guardrails govern the veto. The sample-ratio check runs first because it validates the machinery rather than the hypothesis.

The last item deserves its own paragraph because it is the cheapest bug-catcher in experimentation. A sample ratio mismatch is an arm-size split too improbable under the intended ratio. On 24,000 contacts split 50/50, the standard deviation of the arm count is about 77, so a split of 12,000/12,000 versus 12,150/11,850 is unremarkable — but 12,552/11,448 is more than seven standard deviations out, a probability with a lot of zeros in it. Nobody gets that unlucky. Something is dropping units non-randomly: a bot filter that fires more in one arm, a logging pipeline that loses treatment rows on timeout, an eligibility check evaluated after assignment. Whatever it is, it is correlated with the treatment, which means the arms are no longer comparable and every downstream number is contaminated. Check it first, read metrics second.

Module 03 Power, and what its absence costs

Module 2 left one blank in the design one-pager: how many contacts, for how long. Filling it in is the most consequential arithmetic in experimentation, and it is arithmetic a product manager can do on a whiteboard. This module makes you able to.

The deeper argument here is not about sample size tables. It is that an underpowered experiment is not a weaker version of a good experiment — it is a different instrument with different failure behavior. It misses most real effects, and when it does report a win, that win is systematically too large. A program that runs underpowered tests and ships on their significant results does not merely learn slowly. It learns wrong, confidently, with a chart.

Two ways to be wrong

An experiment can fail in exactly two directions. The Type I / Type II error pair names them: a false alarm, where the change does nothing and you conclude it works, and a miss, where the change works and you conclude nothing. Convention fixes the false-alarm rate at α = 0.05 and then, remarkably, leaves the miss rate β to whatever the sample size happens to produce.

That convention is not neutral. Fixing α while letting β float is a rule that protects the status quo: it makes it hard to wrongly adopt a useless change and says nothing about how often you wrongly abandon a good one. In a regulated control environment this asymmetry is often correct. In a product portfolio deciding which of forty ideas to pursue, it quietly sets your miss rate to a number nobody chose.

Statistical power is the complement of the miss rate: the probability that your experiment produces a significant result if the effect is as large as your minimum detectable effect. It is a property of the design — sample size, variance, MDE, α — not of the result, and it is fully knowable before launch. At 30% power, an experiment on a genuinely working feature comes back “no effect” seven times in ten. Run four such experiments on four real wins and you will ship one.

Two overlapping sampling distributions, null centered at zero and alternative centered at the minimum detectable effect, with the significance threshold and the alpha, beta, and power regions shadedsignificance thresholdnull: the change does nothingalternative: effect = MDE0%−5% (13 min)α — false alarm (5%)β — misspower = 1 − βEvery design choice moves the threshold or the curves; nothing else is available.
Figure 3.1 — The one picture that makes power intuitive. Two distributions of what your experiment might measure: the left curve if the feature does nothing, the right curve if it delivers exactly the MDE. The threshold is where you declare significance. Everything right of it under the null curve is the false-alarm rate α; everything left of it under the alternative curve is the miss rate β; the remainder of the alternative curve is power. Narrowing the curves — more sample — is the only move that shrinks α and β at once.

The calculation

Four quantities are locked in a single relationship: sample size, minimum detectable effect, α, and power. Fix any three and the fourth is determined. There is no fifth lever and no way to buy resolution without paying in one of the others.

For a difference in means between two equal arms, the working formula is:

n per arm  =  2 · σ² · (z_α/2 + z_β)²  /  δ²

Every symbol, in words a skeptic could cross-examine: n is the number of units in each arm. σ is the standard deviation of the metric across units — how spread out resolution times are, the noise you are trying to hear through. δ is the minimum detectable effect in the metric’s own units — the smallest true difference you have designed to catch. zα/2 = 1.96 is the strictness of the significance threshold at a two-sided 5%. zβ = 0.84 is the demand that you catch the effect 80% of the time when it is real. The structure says everything: sample scales with the square of the noise and inversely with the square of the effect you want to see. Halving the MDE quadruples the sample.

Draft Assist, contact-level

σ = 340 minutes. δ = 13 minutes (5% of the 260-minute baseline). α = 0.05 two-sided, power = 0.80.

n = 2 × 340² × (1.96 + 0.84)² / 13²
  = 2 × 115,600 × 7.84 / 169
  ≈ 10,700 contacts per arm

Meridian handles 12,000 contacts a week, split 50/50, so one week yields 6,000 per arm — not enough. Two full weeks yields 12,000 per arm, which clears 10,700 with margin and lands at power 0.84. Verdict: run two full calendar weeks. Whole weeks, because contact mix swings hard by day of week and a run ending on a Wednesday weights the mix.

Now the same question under the design module 2 rejected. Agent-level randomization gives 120 clusters per arm, each contributing roughly 300 contacts over a six-week run. Correlated observations within an agent are discounted by the design effect, 1 + (m − 1) × ICC, where m is contacts per cluster and the intraclass correlation (ICC) — how similar one agent’s contacts are to each other — runs about 0.06 for handle-time metrics.

design effect = 1 + (300 − 1) × 0.06 ≈ 19
effective n   = 36,000 / 19 ≈ 1,900 per arm
MDE floor at that n ≈ 31 minutes ≈ 12%

Run it twelve weeks instead of six and m doubles, but so does the design effect, and effective n converges on clusters ÷ ICC = 120 ÷ 0.06 ≈ 2,000. The floor barely moves. This is the crucial and counterintuitive result: under cluster randomization, duration buys almost nothing once the clusters bind. The agent-level design cannot answer a 5% question at any run length. Only more agents, a larger MDE, or a less variable metric changes the answer.

The load-bearing idea

Power is not a property of how long you ran. It is a property of how many independent units you observed. When units are clustered, adding time adds observations without adding independence, and the experiment plateaus at a resolution set by the number of clusters and their internal similarity.

What underpowering silently does

The familiar cost of low power is missed wins, and it is the least of the three. At 30% power you miss 70% of your genuine improvements — expensive, but at least the failure is visible as a stack of null results and honestly interpretable as “we could not tell.”

The second cost is not visible at all. Significance is a filter, and at low power the filter only passes estimates that noise has inflated. Consider the underpowered case directly: if your design can only reach the threshold when the measured effect is 6% or larger, and the true effect is 2%, then every result you are permitted to call a win is at least three times the truth. This is exaggeration (Type M) error, and it is not a bias in the estimator — an underpowered experiment is unbiased over all its outcomes. It is a bias in the subset you are allowed to notice.

Simulation of the same true two percent effect measured at three sample sizes, showing that only inflated estimates clear the significance threshold at small sample size0%true effect = +2%n = 1,000only this tail is “significant” — it averages ≈ +6%, three times the truthn = 10,000significant results now cluster near the truthn = 100,000nearly every run is significant, and honestred line = significance threshold for that sample size
Figure 3.2 — Underpowered wins are overestimates by construction. A simulation: the same true +2% effect measured three times over, at three sample sizes. The distribution of what you might measure narrows as n grows, and the significance threshold moves in with it. At n = 1,000 the threshold sits far right of the truth, so the only results you are permitted to call wins are the ones noise inflated — they average around +6%. At n = 100,000 the measured effect is the effect. The small study is not a blurry version of the large one; its published results are a biased sample of its own outputs.

The third cost is the one that should change how you read your own program’s track record. Significance answers “how surprising is this data if nothing is happening,” but the question you care about is “given that this came back significant, how likely is it real?” — and that depends on how many of the ideas you test are real in the first place. Suppose 10% of the changes your team ships are genuine improvements, an honest and even generous base rate.

At 80% power: out of 100 tested ideas, 10 are real and you catch 8; 90 are null and you falsely flag 4.5. You report 12.5 wins, of which 4.5 — 36% — are noise. At 20% power: you catch 2 of the 10 real ones, still falsely flag 4.5 of the nulls, and report 6.5 wins of which 4.5 are noise — 69%. Most of that program’s celebrated wins never happened.

Bar chart comparing the composition of significant wins at eighty percent and twenty percent power under a ten percent base rate of real effects8 real4.5 false80% power36% of “wins” are noise2 real4.5 false20% power69% of “wins” are noiseOut of 100 tested ideas, 10 genuinely work (base rate 10%), α = 0.05Height = number of experiments reported as significant wins
Figure 3.3 — The same false positives, a shrinking pile of real ones. Lowering power does not change how many false alarms you generate — α fixes that at 4.5 per 90 null ideas. It changes how many true discoveries sit alongside them. At 20% power the true pile collapses and the false pile does not, so most of what the program celebrates is noise. A team whose experiments are chronically underpowered has a false-discovery problem, not a velocity problem.

The economics of “just run it anyway”

The argument you will actually face is not that power does not matter. It is: “We know it is small, but some evidence beats none — let us just run it and see.” The arithmetic above answers it. An underpowered test consumes traffic and calendar, produces a number that carries the visual authority of evidence, and then contaminates the decision record with a result that is either a miss you will mistake for a refutation or a win you will mistake for a large effect. No evidence at least leaves the question open and the team honest about the state of its knowledge.

There are four legitimate responses when the calculation says you cannot reach power, and each is a real option rather than a consolation:

  • Raise the MDE and say so in writing. If you can only detect 12%, then design for 12% and state plainly that effects below it are invisible to this test. A null result then means “nothing bigger than 12%,” which is a real, bounded finding.
  • Extend the run — but check whether duration actually buys independence. For contact-level randomization it does. For the agent-level design it does not, and the design effect arithmetic is how you prove that to a stakeholder in one slide.
  • Reduce variance rather than adding sample. A more sensitive metric, a trimmed or winsorized version of a heavy-tailed one, or pre-period covariate adjustment (variance reduction techniques of the CUPED family) can cut the required n substantially. These are platform capabilities to ask your data team for; the mechanics are outside this guide.
  • Decide without a test, and record that you did. Some changes are cheap, reversible, and unmeasurable at your traffic. Ship on judgment and label the decision as judgment. This is far more honest than dressing the same judgment in a noise-driven p-value.

One cheap design is legitimate: a pre-declared directional gate — “we will roll back if the point estimate is negative, accepting that this decision rule has roughly a one-in-three error rate at this sample size.” It is legitimate precisely because the error rate is computed and printed in the readout before the run. The same rule invented afterward is the anti-pattern module 4 dismantles.

Anti-pattern — the underpowered test run anyway

Symptom: a design doc with no sample-size calculation, or a run length set by the sprint calendar rather than by δ and σ. Corrective: make the power calculation a required field in the one-pager, and when it fails, force the choice among the four options above in writing. The failure mode is not the missing math; it is that nobody had to say out loud which compromise they were making.

Module 04 Reading results without folklore

The experiment has run. What arrives is a number, an interval, a p-value, and a slide. This module is about reading them without importing the folklore that comes attached — folklore that survives in otherwise sophisticated organizations because the correct statements are one conditional-probability step away from the incorrect ones, and nobody is checking which way the arrow points.

You have a professional advantage here. The most common misreading of a p-value is a fallacy you have already been trained to catch under a different name, in a room with higher stakes. This module leans on that training hard.

What a p-value says, exactly

Here is the sentence, and it repays memorizing: a p-value is the probability of observing data at least this extreme, if the null hypothesis were true. It is a statement about data under an assumption. It is the probability of the evidence given no effect. It is never the probability of no effect given the evidence.

You already own the correction. In a criminal case, an expert testifies that the probability of a random innocent person matching the recovered DNA profile is one in a million. Opposing counsel restates this as a one-in-a-million chance the defendant is innocent. That restatement is the prosecutor’s fallacy, and it is wrong for a reason that has nothing to do with DNA: it inverts the conditional. Converting P(match | innocent) into P(innocent | match) requires a further ingredient the testimony never supplied — how many people were in the candidate pool to begin with. In a city of ten million, roughly ten innocent people match.

Two panels contrasting the probability the test computes with the probability the decision-maker wants, with the inversion arrow struck through and the missing base rate labeledWhat the test computesP( evidence | no effect )“if nothing were happening, data thisextreme would appear 3% of the time”p = 0.03What the decision-maker wantsP( no effect | evidence )“given what we saw, how likely is itthat this feature does nothing?”not computedThe missing ingredient: the base rateHow often do changes like this one actually work? Without that number, the crossing is unavailable.
Figure 4.1 — The conditional runs one way. An experiment computes the probability of the evidence assuming no effect. The decision-maker wants the probability of no effect given the evidence. These are different quantities, and converting between them requires a base rate — how often ideas like this one are real — which the test never supplies. Every p-value misreading in the wild is this arrow, drawn anyway.

The folklore, dismantled item by item. “p = 0.03 means a 3% chance the result is due to chance.” No: it means that if the result were due to chance alone, data this extreme would appear 3% of the time. “p = 0.049 is a win and p = 0.051 is a loss.” No: these are the same evidence, and the threshold is an administrative convention that turns a continuous measure into a binary for the convenience of decision-making. “a very small p means a big effect.” No: p depends on effect size and sample size, so at two million observations a 0.02% change yields a spectacular p-value and no reason to do anything. “p > 0.05 means the feature does nothing.” No, and this one is the most expensive; it is the subject of the next section.

Confidence intervals as the honest summary

A confidence interval is the set of true effect sizes the data cannot rule out. That formulation is doing real work: it turns the result from a verdict into a bounded claim about what remains possible, which is exactly the form a decision needs.

Width is information, and it is the information a p-value discards. Two results can both be “not significant” and mean opposite things. An interval of [−15.9%, +3.5%] says the data is compatible with a large improvement and a modest regression — you know almost nothing. An interval of [−1.5%, +0.7%] says the true effect is within a point of zero in either direction — you know a great deal, namely that this change does not move this metric. The first is ignorance. The second is a finding. Only the second licenses the sentence “this change does nothing,” and no p-value can distinguish them because both report p far above 0.05.

Forest plot of three confidence intervals against a zero line: a clear win, an underpowered ambiguous result, and a tight null, each with its decision verdictno effectClear win−8.1% [−11.4, −4.8]rules out “no effect” and everything smaller than −4.8% → shipUnderpowered−6.2% [−15.9, +3.5]cannot separate a 16% win from a 3% regression → buy informationTight null−0.4% [−1.5, +0.7]rules out anything larger than 1.5% → a real finding−15%−10%−5%0+5%
Figure 4.2 — Three intervals, three different states of knowledge. Read left to right, not by whether the bar touches zero. The clear win excludes zero and everything weaker than −4.8%. The underpowered result spans a large improvement and a real regression — it supports neither shipping nor killing, only the purchase of more information. The tight null also touches zero but rules out anything bigger than 1.5%, which is knowledge, not absence of it. The p-value cannot tell the second from the third.
The load-bearing idea

Absence of evidence is not evidence of absence — unless the interval is tight. A narrow interval around zero is how you convert a non-significant result into a real finding, and it is the only honest route to the sentence “this does nothing.”

Effect sizes over significance

Significance is about surprise. Decisions are about magnitude. A result can be overwhelmingly significant and completely irrelevant: at two million contacts, a 0.02% change in CSAT will produce p = 0.001 and justify precisely nothing, because 0.02% of anything is not worth an operational change. The reverse also happens — a 9% improvement with p = 0.11 at small n is a serious candidate for a larger test, not a rejected hypothesis.

Read the two Draft Assist scenarios as the discipline in action.

Scenario one — the clear win

Two full weeks, contact-level, 12,000 per arm. Mean resolution time falls 21 minutes, −8.1%, 95% CI [−11.4%, −4.8%], p < 0.0001. Guardrails: CSAT flat with an interval excluding any meaningful decline, reopen rate flat, compliance-flag rate flat. Reading: the interval excludes zero and excludes everything weaker than −4.8%, which is comfortably above the 5% MDE, so the effect is not merely real but decision-grade at its lower bound. The guardrail intervals are tight enough to exclude harm rather than merely failing to detect it — a distinction the previous section makes load-bearing. Decision: ship, and forecast on the lower bound, not the point estimate.

Scenario two — the underpowered ambiguity

A follow-up variant, one queue, one week: 1,400 contacts per arm. Resolution time −16.2 minutes, −6.2%, 95% CI [−15.9%, +3.5%], p = 0.21. The deck reads: “directionally positive, recommend ship.” Reading: the data cannot distinguish a 16% improvement from a 3.5% regression. The point estimate is the least informative number on the slide, and — per module 3 — if it had cleared significance at this sample it would have been inflated. Decision: neither ship nor kill. Extend to a powered sample, or redesign the variant. Record that the experiment answered nothing, which is itself worth knowing about the design.

Anti-pattern — “directionally positive”

Symptom: a readout leads with the sign of a point estimate and omits the interval; the recommendation is “trending well, let us ship and monitor.” Why it persists: it feels like appropriate pragmatism, and it is unfalsifiable — every experiment’s point estimate has a direction. Corrective: the decision standard must be a pre-set interval criterion — “ship if the upper bound is below −2%” — declared in the one-pager. Then “directionally positive” either meets the criterion or does not, and the phrase stops doing work it was never entitled to do.

Bayesian and frequentist as working philosophies

This is usually presented as a religious war. It is better understood as two instruments answering two different questions, each with a real cost.

The frequentist offer. Guaranteed long-run error rates: design the procedure and, whatever the truth, you will falsely adopt at most 5% of null changes and detect real ones at your stated power. Nothing needs to be assumed about how likely the effect was beforehand, so nothing about your beliefs is available to be litigated by a stakeholder with an agenda. That property is exactly what a ship/no-ship gate and an audit record want. The price: it answers an oblique question. It tells you how surprising the data would be under a hypothesis nobody believes literally, and it will not tell you the probability that the feature works, because in this framework the effect is a fixed unknown constant and does not have a probability.

The Bayesian offer. A direct answer to the decision-maker’s actual question: given a stated prior and this data, the probability that the effect exceeds the threshold you care about — for instance, an 82% probability the improvement beats 5%. That number composes correctly into an expected-value calculation and updates coherently as data arrives, which makes it the natural instrument for portfolio judgment and continuous monitoring. The price: a prior you must state and defend. This is also its honesty. Everyone brings a prior to a results meeting; the Bayesian writes it down where you can attack it, while the frequentist smuggles it in through the phrase “that seems too good to be true.”

The same weak experimental result updating a skeptical prior and an enthusiastic prior, showing that both posteriors move only modestlySkeptical prior0−6.2%posterior barely movesEnthusiastic prior0−6.2%posterior moves a little furtherdashed = priorgold = likelihood from the dataaccent = posterior
Figure 4.3 — Weak data moves a defensible prior barely, which is the correct amount. The underpowered result (−6.2%, wide interval) is fed to two priors. The skeptical prior barely budges; the enthusiastic one moves further. Neither ends up anywhere that would justify shipping. The divergence between the two panels is small when data is strong and largest when data is weak — which is exactly the region where teams are most tempted to over-read, and exactly why the choice of framework must be made before the numbers arrive.

Working guidance. Use the frequentist gate for ship/no-ship decisions and the compliance record, because guaranteed error rates are what an audit needs and priors are what an audit fights about. Use Bayesian reasoning for portfolio allocation, for continuous harm monitoring, and for any question phrased as “how confident should I be.” Where they disagree materially, the disagreement is a signal that the data is weak and the answer is more data, not a better framework. And distrust anyone who switches frameworks after seeing the result — that is not statistics, it is shopping.

From your other life

The frequentist gate is a burden of proof set before trial: a fixed standard, applied identically regardless of who the defendant is. Bayesian reasoning is how a judge actually updates through a hearing. Both are legitimate; the misconduct is choosing which one governs after the evidence is in.

Module 05 The self-deception catalog

Everything so far assumed good faith and competence. This module assumes both and shows that they are not enough, because each deception in the catalog is a locally reasonable act: checking on your experiment, looking at more than one metric, noticing a pattern in the data. The damage comes from what those acts do to the error rate you thought you had bought — and the error rate does not announce its own inflation. The readout looks identical.

Four entries, one root. Each one moves the goalposts after the ball is in the air: the stopping rule, the metric, the hypothesis. Each has a symptom you can spot in a slide, arithmetic that shows the cost, and a correction that is cheap if committed to in advance and impossible afterward.

Peeking

The experiment is live. The dashboard updates hourly. Someone checks it every morning, and on the day it crosses p < 0.05, the team ships. Every step of this is natural, and the combination destroys the guarantee the 5% threshold was supposed to provide.

The mechanism is easiest to see if you picture the p-value over time. It does not descend smoothly toward the truth; it wanders, because each day’s data adds noise as well as signal. Under a genuinely null experiment, the trajectory random-walks, and a random walk watched at fourteen points has fourteen chances to dip below any line you draw. The 5% guarantee was priced for one look at a pre-specified time. Buy fourteen looks and pay fourteen times — the real false-alarm rate for daily checks over a two-week run lands somewhere around 25%, roughly five times the advertised rate.

A p-value trajectory over a fourteen-day run that dips transiently below the 0.05 threshold on days six through eight and recovers to 0.19 at the pre-registered horizon0.05p = 00.40“it crossed — ship it”pre-registered horizon: p = 0.19day 1day 7day 14Fourteen looks at a wandering statistic: transient significance is the expected outcome, not the surprise.
Figure 5.1 — Where naive stopping ships noise. A p-value does not converge monotonically; it random-walks as data accumulates. Watched daily, a null experiment will dip below 0.05 at some point with far higher probability than 5% — here, on days six through eight — and a team that stops at the dip records a win that the full run would have refused. The threshold has meaning only at the look it was priced for.

The correction costs nothing but discipline, and module 2 already paid for it: the horizon was fixed in the design one-pager before launch. Two additional practices make it enforceable. Report the look count in every readout — “this analysis is the single pre-registered look at day 14” is a sentence that either can or cannot be written. And treat extending a run because the result is not yet significant as the same offense wearing a different hat: it is a stopping rule that depends on the data, which is the definition of the problem.

Anti-pattern — shipping on a peeked p = 0.049

Symptom: the ship decision date is earlier than the horizon in the design doc, and the p-value is just under the line. Mechanism: the threshold was priced for one look; the team took seven, and stopped at the most favorable one. Corrective: read at the horizon, or use a sequential design that has budgeted for the looks in advance. Cost of the corrective: waiting, and occasionally watching a real win sit unshipped for a week.

Sequential testing done right

The prohibition on peeking is unsatisfying, because looking early is genuinely valuable — you want to catch harm, and you want to stop a clear winner sooner. The resolution is that you may peek if you pay, and the payment is arranged in advance.

Group-sequential designs pre-specify the looks — say, at days 4, 7, 10, and 14 — and assign each a stricter threshold, so that the total false-alarm probability across all four still sums to 5%. This is alpha spending: a fixed budget of surprise, allocated across looks rather than spent at one. Early looks demand dramatic evidence (p below roughly 0.001), which is exactly right — stopping after three days should require a result so strong that noise is an implausible explanation. The cost is real: if you spend budget early and do not stop, the final look is slightly stricter than a single-look design would have been, which costs a little power.

The flat naive significance line compared with a group-sequential boundary that demands far stronger evidence at early looks, with the day-seven dip failing to cross the honest boundary0.05naive flat threshold — priced for one lookgroup-sequential boundary — the 5% budget spread across five planned looksday 7: p = 0.028 — under the flat line,nowhere near the honest boundary (0.012)day 1day 7day 14Looking early is allowed. Looking early at the same threshold is not.
Figure 5.2 — Paying for the looks you take. The flat line treats every day as if it were the only day. The sequential boundary demands overwhelming evidence early and relaxes toward the horizon, so that the total chance of a false stop across all planned looks stays at 5%. The day-seven dip that would have triggered a ship under naive reading does not come close to the honest boundary — which is the whole point.

Always-valid inference is the stronger version — confidence sequences that remain correct no matter how often you look — and it is worth naming because it is a platform capability to ask your experimentation team for rather than something to implement yourself. If your platform offers it, continuous monitoring becomes safe by construction and this entire section becomes someone else’s problem.

The asymmetry that matters operationally

Peeking rules protect the ship decision. They do not protect the customer. Monitoring a guardrail for harm and stopping early when it breaches is legitimate, pre-planned, and mandatory — the error you are guarding against there is failing to stop, not stopping too eagerly. Meridian’s compliance-flag guardrail is monitored sequentially with a documented stopping rule from day one, and nobody invokes “no peeking” to keep a harmful variant live.

Multiple comparisons and metric fishing

Run one test at α = 0.05 on a change that does nothing, and you have a 5% chance of a false winner. Run twenty independent tests on the same nothing and the chance that at least one comes back significant is 1 − 0.95²⁰ ≈ 64%. The victory-lap deck — “resolution time was flat, but logins are up 3% and NPS moved two points!” — is not evidence of a hidden benefit. It is the expected output of a null experiment with a wide dashboard.

A grid of twenty metric tiles from a null experiment, one of which reads significant purely by chanceOne experiment that changes nothing, read across twenty metricsresolution timeCSATreopen ratetransfersfirst responsehandle timeescalationsrefund rateloginsp = 0.03NPSchat opt-inself-serve ratequeue waitagent notesmacros usedtickets/agentchurn intentupsell ratesurvey rateAHT variance“Resolution time was flat, but look at logins.”
Figure 5.3 — The lottery, drawn once. Twenty metrics tested at a 5% threshold on an experiment with no true effect: the probability that at least one lights up is 1 − 0.95²⁰ ≈ 64%. The lit tile is not a discovery, it is the arithmetic working as designed. Which tile lights is random; that some tile lights is close to guaranteed.

The corrections, in order of usefulness. First and most important: one pre-registered primary metric. The other nineteen are context, diagnostics, and guardrails — they inform your understanding and can veto, but they cannot promote a null result into a win. This single rule handles most of the damage, and it is free.

Second, for the cases where you genuinely have several co-equal hypotheses, apply a correction to the family: Bonferroni divides the threshold by the number of tests (simple, strict, costs power), or false-discovery-rate control targets the proportion of your declared winners that are false rather than the chance of any error at all (more forgiving, better suited to screening many candidates). The choice is a policy question: are you trying to avoid any false claim, or to keep the false share of your claims tolerable?

Subgroup fishing is the same sin with demographics instead of metrics. “It works for enterprise accounts in EMEA on mobile” multiplies the tests by the number of slices you were willing to cut, and nobody ever counts the slices they looked at and abandoned. If subgroups matter to the decision, name them in the one-pager and treat them as a pre-registered family with a corrected threshold.

HARKing and the pre-registration discipline

The fourth entry is the subtlest, because it involves no statistical error at any single step. HARKing — Hypothesizing After Results are Known — is constructing the hypothesis from the data and then presenting that same data as its confirmation. Every individual move is legal. The composition is circular.

It shows up in a specific narrative shape. The experiment comes back flat. Someone digs, finds that the effect appears concentrated among new customers, and reconstructs a plausible story: new customers have no established workflow, so they adopt the drafted reply rather than fighting it. The story is good. It may even be true. What it cannot be is tested by the data that generated it, because the data was searched over many possible stories and this one was selected for fitting.

From your other life

A theory built to fit the evidence cannot then cite that evidence as independent corroboration. You would take that objection apart in cross without preparation: counsel constructed the timeline after reviewing the documents, then offered the documents as proof of the timeline. The statistical version has better graphics and the same defect.

The honest disposal route is not to suppress the finding — exploratory analysis is valuable and often where the good hypotheses come from. It is to label it correctly and route it to the next experiment. Say: “Exploratory: the effect appears concentrated among accounts under 90 days old. This was not pre-registered and the subgroup was one of eleven examined. We propose a confirmatory test on new accounts with an MDE of 8%, powered at two weeks.” That sentence loses nothing and claims nothing it has not earned.

The instrument that makes all of this enforceable is pre-registration — module 2’s design one-pager, dated and circulated before data exists. Its power is not moral. It is that it makes the difference between a prediction and a postdiction visible, and the deceptions in this catalog all depend on that difference being invisible.

The load-bearing idea

Peeking moves the stopping rule. Multiple comparisons move the metric. Metric fishing moves the outcome definition. HARKing moves the hypothesis. All four are the same act — changing the terms of the test after seeing the result — and all four are prevented by the same artifact, which is a dated document nobody can edit afterward.

Module 06 Metrics that can carry a decision

Module 5 established that the primary metric is the antidote to most self-deception. This module is about how to choose it, what the other numbers are for, and why a guardrail that can be argued away is not a guardrail. The framing is architectural on purpose: a dashboard is not a list of things worth knowing, it is a decision system with roles, authority, and failure modes.

Two of those failure modes deserve their own vocabulary. Metrics get gamed by the people whose work they measure, reliably and quickly. And metrics lie early, in two directions, for reasons that have nothing to do with your feature working.

The taxonomy

Four jobs, and every number on your dashboard holds exactly one of them.

The primary metric is the pre-registered decider. There is exactly one, it is chosen before the run, and the ship decision rides on it alone. Its defining test: would we ship on this metric alone if everything else were unchanged? If the answer is no, it is not your primary.

A guardrail metric is defended, never optimized. It holds veto authority and cannot make the case to ship. Its test: would we block on this alone? A metric nobody would block on is not a guardrail — it is decoration with a serious-sounding name.

A driver metric is a leading indicator of the mechanism. It explains why the primary moved, moves faster than the outcome, and is the first place you look when a result surprises you. Its test: does this explain, rather than decide?

Vanity metrics move with everything and decide nothing — total drafts generated, tokens served, dashboard sessions. They are not evil, they are simply not load-bearing, and the failure mode is that they occupy the top of the deck where a decision-grade number should be.

Quadrant chart placing primary, guardrail, driver, and vanity metrics by decision authority and direction of duty, with the Draft Assist metric set mapped oninforms onlydecidesduty: defendduty: maximizeGuardrailCSAT 78% · reopen 4.1%compliance flags 0.8%rule: would we block on it alone?Primary — exactly onemean resolution timebaseline 260 minrule: would we ship on it alone?Driverdraft open rate · edit distancefirst-reply latencyrule: does it explain, not decide?Vanitydrafts generated · tokens serveddashboard sessionsrule: who games it, how fast?
Figure 6.1 — Four jobs, two axes. Decision authority runs left to right: does this number decide anything, or only inform? Direction of duty runs top to bottom: is it defended against decline, or pushed upward? The primary is the single metric with authority to ship; guardrails have authority only to block; drivers explain the mechanism; vanity metrics claim significance they cannot discharge. A metric that cannot be placed is a metric nobody has decided the purpose of.

Guardrails and the veto

The guardrail’s authority is deliberately asymmetric: it can stop a launch, and it can never authorize one. That asymmetry is the whole design. A metric with symmetric authority is a second primary, and two primaries is one more than the number of things a single decision can ride on.

Scenario three — the guardrail veto

Draft Assist prompt v2 is tuned for a warmer, more confident tone. The readout: mean resolution time down 9.0%, interval clearly favorable, well past the MDE. Guardrails: compliance-flag rate 0.8% → 1.9% (+1.1pp, 95% CI [+0.7, +1.5pp]); reopen rate 4.1% → 5.3% (+1.2pp, 95% CI [+0.5, +1.9pp]); CSAT flat.

Mechanism diagnosis. These are not two independent problems and a win. They are one behavior seen three ways. The v2 drafts are faster because they are confidently wrong: they commit to a refund timeline or a policy reading without hedging, agents send them with less editing, the contact closes quickly — and then QA flags the misstatement and the customer writes back. Speed, flags, and reopens are the same phenomenon measured at three points in time.

Decision: veto. A 1.1-point rise in replies that misstate refund terms is a control failure with regulatory exposure; there is no resolution-time number that buys it. The information dividend: the breach identified the mechanism precisely, which tells you what v3 must do — gate drafts behind a compliance classifier that blocks or flags any draft asserting a refund amount, timeline, or policy exception, then re-test. The failed experiment paid for the next design.

Two disciplines make a veto real. First, guardrail thresholds are set before the run, in the one-pager, as pre-specified breach conditions — “any increase in compliance-flag rate whose interval excludes zero halts rollout.” A threshold negotiated after the data arrives is a negotiation, not a control. Second, guardrail intervals must be read for whether they exclude harm, not for whether they reached significance: a flat point estimate with an interval spanning [−2pp, +3pp] on compliance flags has not cleared anything, it has failed to look.

From your other life

Guardrails are compensating controls, and a breached guardrail is a control finding, not a footnote. You already know what happens to a program that treats control findings as inputs to a cost-benefit conversation with the business owner who wants to ship — the control stops being a control the first time it is overruled, and everyone learns that in one meeting.

Drivers, outcomes, and the proxy chain

You randomize on what moves in weeks; you care about what moves in quarters. The bridge between them is a chain of proxies, and its links are assumptions of wildly different quality that a dashboard renders as identical rectangles.

The Draft Assist proxy chain from draft adoption through resolution time to cost per contact and retention, with the evidence status of each link marked and the weakest link flaggedDraft adoptiondriverResolution timeprimaryCost per contactand CSAToutcomeRetentionthe thing we wantrandomizedmeasured in this testarithmeticminutes × loaded rateassumed — the weak linkobservational only, confoundedEvery arrow is a claim with its own evidence status. The dashboard draws them all the same width.
Figure 6.2 — The chain, with each link’s evidence status marked. Adoption to resolution time is randomized in this experiment. Resolution time to cost is arithmetic. Resolution time to retention is an assumption supported only by confounded observational data — the same shape as module 1’s app-adoption claim. The chain is exactly as strong as its weakest arrow, and naming which arrow that is prevents a randomized result from being quoted as evidence for an unrandomized conclusion.

The characteristic failure is promoting a driver to primary because it is convenient. Draft-open rate moves within hours, has tiny variance, and produces beautiful charts — and measuring it tells you that agents open drafts, not that anything got better. Optimizing a driver directly is how you end up with a feature that maximizes its own usage. The primary must be the nearest metric downstream that would justify shipping on its own.

Goodhart dynamics

Goodhart’s law: when a measure becomes a target, it ceases to be a good measure. The mechanism is not mysterious. People whose compensation, ranking, or standing depends on a number will move that number, and moving the number is almost always cheaper than moving the thing the number was a proxy for.

Meridian has already run this experiment without meaning to. When handle time became a floor-level target, it improved 15% in a quarter — and reopen rate climbed, because the fastest way to close a ticket is to close it before the customer is finished. The handle-time improvement was real and the underlying performance was worse.

The design pattern that survives this is metric–guardrail pairs: for every metric you make a target, name the cheapest way to game it and instrument the tripwire that gaming trips. Speed pairs with reopen rate and CSAT. Volume pairs with quality sampling. Automation rate pairs with escalation rate. Cost per contact pairs with repeat-contact rate. The pair is not a nicety; a target without its tripwire is an instruction to game.

From your other life

Red-team the metric before it ships. Take the adversarial posture you would take toward an authorization model and point it at the incentive: who is measured on this, what is the cheapest action that moves it, does that action serve the customer, and what would detect it? Treat “nobody would do that” as the answer that has never once been correct about an incentive.

Metrics that lie early

Two transients distort early readings in opposite directions, and both resolve toward the truth if you let the experiment run.

Novelty: people engage with a new thing because it is new. Usage spikes, the metric it drives spikes, and the effect decays toward truth from above. Primacy: a change disrupts practiced workflows, so performance dips while people relearn, and the effect recovers toward truth from below. Draft Assist is a primacy-shaped feature: agents who have written replies from scratch for four years must learn a different motion — read, judge, edit — and they are slower at it before they are faster.

Two treatment-effect trajectories over six weeks, one showing a novelty spike decaying toward a modest truth and one showing a primacy dip recovering to a positive effect, with the week-one reading marked0+12%−8%novelty: spikes, then decays toward the truthprimacy: dips while people relearn, then recoversthe week-1 readout: overstates one feature, condemns the otherwk 1wk 3wk 6Compare the first week against the last before believing either.
Figure 6.3 — Two transients, opposite signs, same cure. A novelty-shaped effect starts large and settles small; a primacy-shaped effect starts negative and settles positive. Read at week one, the first feature looks like a triumph and the second like a failure, and both readings are artifacts of when you looked. Running through the transient and comparing the first week of exposure against the last is the diagnostic that separates a decaying artifact from a stable effect.

The detection method is a shape, not a test: plot the treatment effect by week of exposure rather than calendar week, so that users entering later are aligned by their own experience rather than by the date. If the first week and the last week disagree materially, the run was not long enough, and the honest readout says which direction the effect is still moving.

Anti-pattern — shipping on the transient

Symptom: a decision made from the first days of exposure, usually justified by early significance. Corrective: set run length from the transient as well as from power — long enough for the learning curve to flatten — and report the week-one versus week-final comparison in the readout as a standard field. Note the asymmetry of cost: shipping on novelty wastes engineering on an effect that will evaporate; killing on primacy discards a feature that works, and you will never learn that you did.

Module 07 When you can’t randomize

Randomization is not always available, and the reasons are usually good ones: a regulator or works council forbids differential treatment, a fairness constraint makes withholding a benefit indefensible, a platform change is physically all-or-nothing, or the contamination between arms is so severe that a randomized comparison would measure the mixture rather than the treatment.

What remains is a family of designs that reconstruct the counterfactual from structure rather than from a coin. They are not consolation prizes — a well-argued difference-in-differences beats a badly contaminated experiment — but they shift the burden. In a randomized trial, the identifying assumption is enforced by the mechanism. In a quasi-experiment, it must be argued to a skeptic, and that argument, not the point estimate, is the deliverable.

The fallback hierarchy

When the coin flip is off the table, the options are not equally good, and the ordering principle is simple: how much of the identification is guaranteed by design versus supplied by your argument.

DesignWhere the counterfactual comes fromWhat you must argue
Randomized experimentThe coin flipOnly that the mechanism worked (SRM, A/A)
Natural experimentAn external shock that assigned treatment for reasons unrelated to the outcomeThat the shock was genuinely unrelated
Designed quasi-experiment (DiD, RD)Structure: a shared time trend, or an arbitrary thresholdA specific, stateable assumption — parallel trends, no manipulation
Adjusted observationalStatistical adjustment on measured covariatesThat you measured every confounder that matters
Raw correlationNowhereEverything

The fourth row is where most business analysis lives and why most business analysis is weak: “we measured every confounder that matters” is an unfalsifiable claim about your own imagination, and module 1 showed why more control variables do not strengthen it. The third row is different in kind. Parallel trends and no-manipulation are specific claims with observable implications, which means they can be tested against evidence, attacked by a skeptic, and defended or abandoned on the merits.

From your other life

This is a burden-of-proof shift, and you should treat it procedurally rather than as a matter of degree. Under randomization, the mechanism carries the burden. Under a quasi-experiment, you carry it — which means the deliverable is not a number but a memo: here is the comparison, here is the assumption it rests on, here is the evidence for the assumption, and here is the single fact that would defeat it.

Difference-in-differences

Meridian cannot randomize Draft Assist across its whole support organization, because its EU works council must approve any tool that monitors or augments agent work, and that approval will not arrive before Q3. Legal forces a staggered rollout: North America gets the feature in March, the EU does not. This is a constraint, and it is also a design.

The naive read is the before/after on North America: 262 → 241, so the feature saved 21 minutes. Module 1 already refused that, because it substitutes February for the counterfactual. But now you have something February never gave you — a group living through the same calendar without the treatment. The EU went 259 → 254, a 5-minute improvement with no feature at all: tooling upgrades, seasonal mix, a maturing support playbook, the ordinary secular drift that would have lifted North America too.

NA:  262 → 241   change = −21 min
EU:  259 → 254   change = −5 min
difference-in-differences = (−21) − (−5) = −16 min  (−6.1%)

The second difference is the whole trick: it subtracts whatever both regions experienced in common. Whatever the shared drift was — seasonality, the new knowledge base, a quieter billing cycle — it hit both and cancels.

Difference-in-differences: North America and EU resolution-time lines across pre and post periods, with the counterfactual North America line drawn on the EU trend and the estimate shown as the resulting gap262241Draft Assist launches in NA (March)North America (treated)EU (comparison — no feature)NA counterfactual: the EU trend applied to NADiD = −16 min (−6.1%)8 pre-period weeks: the NA–EU gap holds within ±3 minutespost-periodThe naive before/after would have claimed −21. The EU’s −5 shows how much of that was drift.
Figure 7.1 — The second difference subtracts the shared trend. North America falls 21 minutes; the untreated EU falls 5 over the same calendar, revealing a secular improvement that would have reached North America anyway. Projecting the EU trend onto the North American pre-period level produces the dashed counterfactual, and the estimate is the vertical gap between where North America actually landed and where the shared trend says it would have. The design lives or dies on whether that dashed line is credible.

Which brings us to the assumption, stated so a skeptic can swing at it: absent Draft Assist, the gap between North American and EU resolution time would have stayed roughly constant. That is parallel trends. It is not the claim that the regions are similar — they are not; the EU runs three minutes faster at baseline and that is fine. It is the claim that whatever makes them differ is stable across the period.

You argue it with evidence. Meridian has eight pre-period weeks in which the NA–EU gap held within ±3 minutes against a 16-minute estimated effect — a gap that stable does not suddenly move five times its historical variation by coincidence. You also argue it by exclusion: what could break it? Any change affecting one region and not the other during the window. An EU-only migration to a new ticketing system. A North American staffing surge. A regional product launch that changed contact mix. A holiday calendar that falls differently. Each of these is checkable, and the memo should say which ones you checked.

What kills a difference-in-differences

Not a violation of similarity — a violation of parallelism. And note what does not repair it: adding control variables after the fact. If the EU migrated ticketing systems mid-period, no covariate adjustment recovers the counterfactual, because the comparison group stopped being a valid stand-in for the treated group’s trend. The honest response is to shorten the window to exclude the shock, find a cleaner comparison region, or abandon the claim.

Regression discontinuity

Your organization already enforces arbitrary thresholds — spend tiers, tenure cutoffs, score bands, SLA classes — and every one of them is a small experiment nobody ran on purpose.

Meridian assigns a dedicated customer success manager to accounts above $10,000 in annual spend. The design intuition: an account at $9,950 and an account at $10,050 are, in every respect that matters, the same kind of account. Nothing about the business changes across a hundred dollars of spend — except the rule, which fires on one and not the other. In a narrow window around the cutoff, CSM assignment is as good as randomized, so comparing retention just above the line against just below estimates the CSM effect for accounts near the threshold.

Regression discontinuity: retention plotted against annual spend with a jump at the ten thousand dollar cutoff, plus a small panel showing the density pile-up that would invalidate the design92%78%annual spend →$10,000 — CSM assignment rulethe estimatethe check that saves youaccount densitya pile-up at $10,001 meanssomeone is steering accounts
Figure 7.2 — A jump at an arbitrary line. Retention rises smoothly with spend, so the level difference between a $5,000 and a $30,000 account tells you nothing causal. The discontinuity does: at exactly the point where the rule fires and nothing else changes, retention jumps, and that jump is the CSM effect for accounts near the threshold. The right-hand panel is the mandatory diagnostic — if accounts pile up just above the cutoff, someone is choosing which side to land on, and the two groups are no longer comparable.

Two assumptions, both stateable in a sentence. No precise manipulation: nobody can choose which side of the line they land on. This fails the moment a sales rep learns that nudging an account to $10,001 wins it CSM coverage, and it fails visibly — the density of accounts spikes just above the threshold, which is why the histogram check is mandatory rather than optional. Nothing else changes at the same threshold: if $10,000 also triggers a different support tier, priority routing, and a quarterly business review, the jump measures the whole bundle, and no amount of data separates the CSM from its co-travelers.

The scope limit is important and often forgotten: regression discontinuity estimates the effect near the cutoff. It tells you what a CSM does for a $10,000 account. It says nothing about what a CSM would do for a $200,000 account, where the relationship may be entirely different — and that extrapolation is the most common way a valid estimate gets misused.

Honest observational claims

The final defense is linguistic, and it is the one you exercise most often, because most claims that cross your desk will never be upgraded to any design at all. Each design licenses a specific sentence, and the discipline is refusing to speak above your evidence.

Evidence you haveThe sentence you may sayThe sentence you may not
Raw correlation“X and Y move together in our data.”Anything with a causal verb
Correlation + adjustment“The association persists after adjusting for size, tenure, and queue mix.”“Controlling for confounders, X causes Y”
Difference-in-differences“Relative to a comparison region on a parallel pre-trend, X was followed by a 16-minute improvement; the estimate holds if no region-specific shock occurred.”“X caused a 16-minute improvement” without the conditional
Regression discontinuity“For accounts near the $10,000 threshold, the rule raises retention by N points.”Any claim about accounts far from the cutoff
Randomized experiment“X causes a 21-minute reduction, 95% CI [.., ..], in this population over this window.”Claims outside the tested population or window

Two habits round this out. The first is a sensitivity question you can ask without any math: how large would an unmeasured confounder have to be to erase this estimate? If a modest, entirely plausible difference in customer engagement would account for the whole gap, the claim is fragile regardless of its p-value. If it would take a confounder larger than any variable you have ever measured, the claim is robust — and that is an argument a skeptic can engage with rather than a defensive assertion that you controlled for things.

The second is knowing when observational evidence is genuinely enough. It often is: for decisions that are cheap and reversible, for effects far too large for any plausible confounder to manufacture, and for decisions you would make anyway where the analysis only calibrates expectations. The cost of over-demanding rigor is real — a team that will not act without an RCT will be beaten by one that acts well on good observational evidence. The discipline is not refusing to decide. It is refusing to misdescribe what the decision rests on.

Module 08 Experimenting on LLM features

LLM features attract two opposite errors. The first is that they are so novel that ordinary experimental discipline does not apply — usually expressed as shipping on eval scores and vibes. The second is that they are identical to any other feature and need no special handling, which walks straight into version drift and per-response variance.

Both are wrong in the same way: what changes is the variance structure and the confounders, not the logic. The counterfactual argument, the power arithmetic, and the entire self-deception catalog apply unmodified. This module covers the three things that genuinely differ — the unit calculus, the offline-to-online linkage, and the version confounders — using Draft Assist’s prompt-v2 test as the worked case.

What is actually different

Four things change, and each has a design consequence rather than merely being interesting.

Nondeterminism. The same contact submitted twice yields different drafts. This is not a bug to be eliminated — temperature is often load-bearing for quality — but it means a single response is a draw from a distribution, and per-response quality metrics carry more variance than a deterministic feature’s would. It also means an A/A test on an LLM feature is genuinely informative: two arms running the identical prompt will differ, and knowing how much they differ is your noise floor.

Version drift. The provider ships a model update mid-run. A teammate improves the system prompt on Tuesday because it obviously needed it. The retrieval corpus gets re-indexed. Each of these changes the treatment while the experiment is measuring it, and none of them appears in your experiment platform. This is the single most common way an LLM experiment silently becomes uninterpretable.

Per-response variance dwarfs per-user variance. Draft quality varies enormously between contacts — a password reset drafts perfectly, a multi-invoice billing dispute does not — and that spread is typically much larger than the spread between agents. That is good news for power on response-level metrics and irrelevant to power on agent-level outcomes, which is a distinction the next section makes operational.

Cost is a first-class metric. Tokens are cost of goods sold, and a variant that improves resolution time by 6% while tripling inference cost may be worse than the control. Cost per contact belongs on the readout as a standing guardrail, not as a footnote after launch.

What has not changed

The counterfactual logic of module 1, the unit and interference reasoning of module 2, the power arithmetic of module 3, the interval discipline of module 4, the entire self-deception catalog of module 5, and the metric architecture of module 6 all apply without modification. LLM features get no exemption; they get extra confounders.

The unit decision, revisited

Module 2’s rule holds: match the unit of randomization to the behavioral footprint of the treatment. LLM features make this sharper, because the same product can host changes with radically different footprints.

Response-level. Each individual draft is independently generated by prompt v1 or v2. Correct when the change is invisible to workflow — prompt wording, model swap, retrieval tweak, temperature. The agent cannot tell which variant produced the draft in front of them, so there is nothing to learn and nothing to contaminate. Maximum power: Meridian generates 12,000 drafts a week, and per-response metrics like edit distance and draft acceptance stabilize within days.

Contact- or session-level. Correct when the treatment shapes a whole interaction — a multi-turn assistant, a change that affects the second reply given the first. Randomize the conversation, not the turn, or the arms interleave inside a single customer experience.

Agent-level. Required when the treatment changes workflow or trains behavior. The original Draft Assist launch was agent-visible: the feature’s existence changes how an agent works, and an agent who has drafts on half their contacts learns differently than one who has them always or never. Its prompt version, by contrast, is invisible — which is why the v2 test runs response-level on a fraction of traffic while the launch test could not.

Grid comparing response, contact, and agent randomization for LLM features across the question each answers, available power at Meridian traffic, and contamination riskUnitThe question it can answerPower at 12,000/weekContaminationResponseprompt v2 testIs this draft better than that draft?prompt wording · model swap · retrievalvery high — days, not weeksnone if the changeis invisible to the agentContactDoes the whole interaction go better?multi-turn behavior · first-reply effectshigh — two weeks at 5% MDEagent-learningspillover, toward nullAgentoriginal launchDoes having the feature change howsomeone works?low — effective n caps ≈ 2,000clean exposureThe mixture trap: randomizing responses when the outcome lives at the agent level — both arms are the mixture, and neither is a pure condition.
Figure 8.1 — Match the unit to the treatment’s behavioral footprint. A change the agent cannot perceive can be randomized at the finest grain, which buys enormous power. A change that alters how someone works must be randomized at the level of the person who works, or both arms end up experiencing a blend of conditions and the comparison measures nothing that would exist after launch.

The trap in the dashed band deserves naming because it is easy to walk into while optimizing for power. If you randomize drafts on and off at the response level and measure resolution time, every agent experiences a 50/50 mixture. They adapt to the mixture — hovering, waiting to see whether a draft appears, developing a hybrid workflow. Neither arm reflects the world where the feature is always on or always off, and the measured difference is a within-mixture contrast that will not survive launch.

Linking offline evals to online metrics

An offline eval is a versioned test suite for model behavior: fixed inputs, scored outputs, run in CI before anything reaches a customer. Guide Nº 03 covers how to build one. The question here is what an eval score means in the evidence hierarchy this guide has built.

The answer is precise: an eval is a driver metric. It sits at the far left of module 6’s proxy chain, it moves in minutes rather than weeks, it explains rather than decides, and — like every driver — it earns its position only by predicting the outcome it proxies for. The division of labor is the familiar one from software: evals are the unit tests, cheap and fast and run constantly; the online experiment is the integration test against reality, expensive and slow and authoritative.

A loop from the offline eval suite through the online experiment and back through recalibration, with the ship-on-evals-alone shortcut drawn as a dashed bypassOffline — cheap, fast, no customersOnline — expensive, slow, authoritativeprompt candidateseval suite gatedriver metriconline experimentprimary + guardrailsship decisionrecalibrate the eval against what actually happened“the eval improved — ship it”the bypass that skips realityWhen eval and experiment disagree, reality wins and the eval gets retuned. Never the reverse.
Figure 8.2 — The loop, and the shortcut that breaks it. Candidates are screened offline, survivors are measured online, and the online result feeds back to recalibrate the eval — which is what keeps the cheap filter predictive over time. The dashed bypass ships on eval improvement alone: it is fast, it feels rigorous because a number went up, and it substitutes a proxy for the outcome the proxy was supposed to predict.

The linkage discipline is concrete. Across your last several launches, plot the eval delta against the measured online effect. If the relationship is real, the eval has earned its gate authority and you can screen aggressively with it. If the scatter is a cloud — and it often is, for evals assembled from cases that were easy to write rather than cases that predict outcomes — then the eval is measuring what someone thought mattered, and it needs rebuilding against cases drawn from real traffic.

When they disagree on a specific launch, the resolution is not symmetric. Reality wins: the experiment measured what customers and agents actually did, and the eval measured a proxy. The eval gets recalibrated — add the cases the online result revealed, re-weight the scoring — and it keeps its job as the cheap pre-filter. What does not happen is the experiment being re-run until it agrees with the eval.

Anti-pattern — shipping on eval improvement alone

Symptom: a launch recommendation whose evidence section contains only offline scores, usually for a change described as “just a prompt tweak.” Corrective: require an online result for anything that reaches a customer, and set the bar by blast radius — a response-level prompt test costs days, which is affordable for nearly every change worth making. Where the change is genuinely too small to test, say so explicitly and record it as an untested ship.

Reading LLM experiments

Two practical matters close the module: how long these tests need to run, and what must be frozen while they do.

Duration follows the metric, not the feature. Per-response quality metrics — draft acceptance rate, edit distance between the draft and the sent reply, hallucination-report rate — have large per-response variance but enormous sample counts, and they stabilize in days. Agent-behavior and customer metrics — resolution time, CSAT, reopens — obey module 3’s arithmetic unchanged, which for Meridian means two full weeks at contact level. A prompt-v2 test can therefore report acceptance and edit distance after three days and must still wait for resolution time, and a readout that presents the fast metrics as if they settled the slow question is the driver-as-primary error in a new costume.

Freeze the treatment. Three things must be pinned for the run’s duration and stated in the one-pager: the model version (pin an explicit snapshot identifier, not a floating alias; a provider upgrade mid-run is an uncontrolled treatment applied to both arms at once), the prompt and its template (no improvements during the run, however obvious — an edit on day 4 means days 1–3 and 5–14 tested different things), and the retrieval corpus version if the feature retrieves. Log the version identifiers alongside every response so that if something moves anyway, you can segment the analysis at the boundary instead of discarding the run.

Guardrails for a generation feature extend module 6’s set with the failure modes specific to generated text: compliance-flag rate (already Meridian’s sharpest instrument, since it detects exactly the confident-and-wrong failure), hallucination-report rate from agents, escalation rate, and cost per contact in tokens. Each is defended, none is optimized, and each has a pre-specified breach condition.

Cross-reference

Guide Nº 03 builds the eval suite this module points online. The relationship runs both directions: an eval that has never been checked against an online result is an untested proxy, and an experiment program with no eval layer pays full experimental cost for changes that a three-minute offline run would have rejected.

Module 09 The unfoolable PM

Eight modules of machinery reduce to one capability: reading a results readout and knowing, within a few minutes, whether it can carry the decision someone is about to make with it. This module is the reference layer — the catalog in one table, the argument that null results are assets, and the ten questions that turn a slide into testimony.

None of this requires you to run the analysis. It requires you to know which questions are lethal and why, which is the skill your first career was built on and the reason this guide has leaned on it throughout.

The catalog, consolidated

Six patterns account for most of the bad decisions made from good data. Each is stated as a symptom you can spot without access to the underlying analysis, because that is the situation you will usually be in.

Anti-patternSymptom in the wildCorrectiveTaught in
Shipping on a peeked p = 0.049The decision date precedes the documented horizon; the p-value sits just under the lineRead at the horizon, or specify a group-sequential boundary before launch; report the look countModule 5
“Directionally positive” as a decision standardThe recommendation cites the sign of a point estimate; no interval appears on the slideA pre-registered interval criterion — ship only if the upper bound clears a stated thresholdModule 4
Twenty metrics and a victory lap on the two that movedThe primary is flat; secondary metrics carry the recommendationOne pre-registered primary decides; corrected thresholds for genuine families; count the subgroups examinedModule 5
Randomizing agents, analyzing contactsThe design says one unit and the sample size says another, orders of magnitude apartAnalyze at the randomization unit or cluster the standard errors, and say whichModule 2
Ignoring interference on a shared queueThe control arm moves during the run; nobody asks whyEnumerate shared resources; monitor control-arm conditions; isolate routing or reason out the bias directionModule 2
Treating a failed experiment as a failed featureA null result ends a program; nobody states what the interval ruled outRead the width: a tight null is a finding, a wide null is a request for more informationThis module

Notice the shape of the correctives. Five of the six are artifacts — a document, a threshold, a field in a template — rather than acts of vigilance. That is deliberate. Vigilance does not survive a quarter under deadline pressure; a required field in a readout template does.

Failed experiments are information

A null result with a tight interval is a finding, and treating it as a failure is the most expensive habit in this guide because it corrupts what a team is willing to test.

Draft Assist v3 returns −0.4%, 95% CI [−1.5%, +0.7%]. That is not “we learned nothing.” It is: this variant does not move resolution time by more than 1.5% in either direction, and we will not be revisiting that hypothesis. A roadmap myth has been killed cheaply and permanently. Compare the alternative — shipping it on a wide interval and spending two quarters attributing unrelated fluctuations to it.

The portfolio frame follows directly from module 3’s arithmetic. If roughly one idea in ten genuinely works, a well-run program has a low win rate by construction. A team reporting that eight of nine experiments were wins is not outperforming the base rate by a factor of eight; it is exhibiting the catalog — peeking, metric flexibility, post-hoc subgroups — or testing changes so certain they did not need an experiment. Either way the number is a symptom, and the correct response to it is curiosity about their protocol, not envy.

The distinction that closes the thesis

The experiment refuted this implementation’s effect on this metric, in this population, over this window. It did not refute the hypothesis, the mechanism, or the next variant. Popper’s asymmetry cuts both ways: a refutation is decisive about exactly what was tested and silent about everything else. Teams that forget the first half ship on noise; teams that forget the second half kill good ideas on one bad implementation.

Two practices operationalize this. Score experiments on information per week rather than on wins: a tight null that closes a question is worth more than a wide win that reopens three. And make the retro question “what does this rule out?” rather than “what went wrong?” — the first produces a durable record, the second produces defensiveness and no record at all.

The cross-examination

Ten questions, in deposition order — general foundation first, then the specific attacks, so that a readout which fails early does not waste the room’s time on interpretation.

  1. What was the pre-registered primary metric, and can I see the document? (Module 2, 5 — if there is no dated artifact, everything downstream is a story told after the fact.)
  2. Was the horizon fixed in advance, and how many times was the result looked at? (Module 5 — the threshold has meaning only at the look it was priced for.)
  3. How many metrics and how many subgroups were examined, including the ones not shown? (Module 5 — twenty metrics produce a winner about two times in three.)
  4. Does the analysis unit match the randomization unit? (Module 2 — a mismatch fabricates significance and voids the arithmetic itself.)
  5. Could the arms have touched each other? (Module 2 — shared queues, shared inventory, shared social channels; and which direction does the leak bias the result?)
  6. Is the measurement window past the novelty and primacy transient, and do first-week and last-week effects agree? (Module 6.)
  7. What does the interval rule out — and what does it fail to rule out? (Module 4 — the width, not the p-value, is the state of knowledge.)
  8. Did any guardrail move, and does its interval exclude harm or merely fail to detect it? (Module 6 — a wide flat guardrail has not cleared anything.)
  9. Was the assignment split checked before the metrics were read? (Module 2 — a sample ratio mismatch invalidates everything downstream.)
  10. Who computed this, and what happens to them if it says no? (All modules — not an accusation; the incentive tells you which of the previous nine to press hardest.)
A results readout passing through five sequential checks, with each failure routing to its named deception and corrective, and survivors exiting to decision-grade evidenceresultsreadoutprimarypre-registered?horizon fixed,looks counted?analysis unit= random. unit?arms isolated,past transient?interval read,guardrails clear?metric fishingm5 · pre-register onepeekingm5 · fix the horizonfabricated significancem2 · cluster the errorsinterference / transientm2, m6 · isolate, run longer“directionally positive”m4 · read the widthdecision-grade evidencenow argue about the decisionFailures upstream void the checks downstream: a fabricated standard error makes the interval meaningless,so ask the questions in order and stop at the first one that fails.
Figure 9.1 — The gauntlet. A readout enters at the left and passes five sequential checks; each failure routes to a named deception with its corrective and its home module. The ordering is load-bearing: an upstream failure invalidates every downstream check, so a unit mismatch makes the interval discussion pointless. Only a readout that survives all five is evidence you can argue a decision from — and then the argument is about the decision, not about the number.
From your other life

These are leading questions for a friendly witness. You are not trying to destroy the analyst — you want them to walk in already knowing the ten and having answered them in the document. The goal of a cross-examination checklist that everyone has read is that it never has to be conducted.

Concept index

Counterfactual
The unobserved outcome the same units would have had without the treatment — what every causal method tries to reconstruct.
Confounder
A variable that drives both the treatment and the outcome, manufacturing correlation without causation.
Backdoor path
The route correlation travels through a confounder’s two edges instead of through the causal edge.
Selection bias
Distortion that arises because who ends up in your data is itself an outcome, not a random draw.
Survivorship bias
Analysis restricted to units that survived to be measured, hiding the failures that exited.
Randomization
Assignment by coin flip, which severs every backdoor path — measured or not — at once.
Unit of randomization
The entity the coin flip is applied to (contact, agent, queue); the experiment’s load-bearing design choice.
Interference (SUTVA violation)
One unit’s treatment leaking onto another unit’s outcome, for example through a shared queue.
Intention-to-treat
Analyzing units by what they were assigned, not what they used — preserving what randomization bought.
Sample ratio mismatch
An arm-size split too improbable under the intended ratio; the smoke detector for broken assignment.
A/A test
An experiment with no treatment difference, run to confirm the machinery finds nothing.
Minimum detectable effect (MDE)
The smallest true effect the experiment is designed to detect reliably; chosen from business stakes, not hope.
Statistical power
The probability of detecting the MDE if it is real; one minus the miss rate.
Type I / Type II error
False alarm (calling a null effect real) / miss (calling a real effect nothing).
Exaggeration (Type M) error
The systematic overestimate in significant results from underpowered tests — the significance filter at work.
Intraclass correlation (ICC)
The within-cluster similarity of outcomes that shrinks a cluster-randomized sample’s effective size.
Design effect
The multiplier by which clustering inflates the sample you need, driven by cluster size and ICC.
p-value
The probability of data at least this extreme if the null were true — the chance of the evidence given no effect, never the reverse.
Confidence interval
The set of true effects the data cannot rule out; its width is the honest measure of what you know.
Practical significance
Whether the effect is big enough to matter, a question the p-value does not address.
Prosecutor’s fallacy
Reading the probability of the evidence given innocence as the probability of innocence given the evidence.
Prior / posterior
What you believed about the effect before the data, and what you should believe after — Bayes’ stated assumptions.
Peeking
Repeatedly checking results and stopping at transient significance, inflating the false-positive rate several-fold.
Alpha spending
Distributing the 5% budget of surprise across planned looks so early stopping stays honest.
Multiple comparisons
Testing many metrics or subgroups at once, guaranteeing spurious winners by volume.
HARKing
Hypothesizing After Results are Known — building the story from the data, then citing the data as its proof.
Pre-registration
Committing the primary metric, horizon, and analysis in writing before the data exists.
Primary metric
The single pre-registered metric the ship decision rides on.
Guardrail metric
A metric with veto power that is defended, never optimized.
Driver metric
A leading indicator of the mechanism — explains why, but cannot justify shipping alone.
Goodhart’s law
When a measure becomes a target, it stops being a good measure — because people paid on it attack it.
Novelty / primacy effect
Opposite early-reading transients: inflated effects from curiosity, deflated effects from disrupted habit.
Difference-in-differences
Comparing treated and untreated groups’ changes over time to subtract the shared trend.
Parallel trends
The load-bearing assumption of difference-in-differences: absent treatment, the groups’ gap would have held steady.
Regression discontinuity
Treating units just either side of an arbitrary cutoff as as-good-as-randomized.
Offline eval
A versioned pre-ship test suite for model behavior — the unit test whose predictive link to online results must be earned.

This is the complete text of the course. With JavaScript enabled, this same page runs the interactive edition — a self-diagnostic that reorders the syllabus around your gaps, knowledge checks, applied worksheets with model answers, and a spaced-repetition review queue — with progress saved locally in your browser.