Beyond the Dashboard · Nº 10
A field guide to experimentation & statistical judgment
Every causal claim your organization makes is a comparison against a world that does not exist: the same customers, the same quarter, without the change. You never observe that world. Everything in this guide — randomization, power, intervals, difference-in-differences — is a strategy for constructing a credible stand-in for it, and every failure of judgment in the field is a bad stand-in accepted without argument.
This module is about the three ways a bad stand-in gets accepted. Each has a name, a mechanism, and a tell you can spot in a slide. You are already trained for this work: a correlational claim is circumstantial evidence, and your instinct to ask who selected the exhibits, what else could explain them, and who is missing from the room is exactly the right instinct. What follows gives that instinct a vocabulary precise enough to argue with an analyst.
Meridian, a subscription-billing company, ships Draft Assist — an LLM feature that pre-drafts replies for its 240 support agents. In March, mean resolution time falls from 262 minutes to 241. The deck says the feature saved 21 minutes per contact. It says no such thing yet, because the claim being made is about a comparison, and only half of the comparison has been observed.
The counterfactual is the outcome the same units would have had, over the same period, without the treatment. It is unobservable by construction — you cannot both ship and not ship to the same contacts in the same week. This is the fundamental problem of causal inference, and it has exactly one class of solution: find or build a group whose observed outcome is a credible substitute for the missing branch.
Notice what the March dashboard actually substituted for the dashed branch: February. That substitution assumes nothing else changed between the two months — no seasonal dip in dispute volume, no staffing change, no queue-routing update, no post-holiday billing surge working its way out of the system. Before/after comparison is not a weak experiment. It is a strong assumption wearing a chart.
“Compared to what?” is the first cross-examination question, and it has only three honest answers: a randomized control group, a designed quasi-experimental comparison whose assumptions you are prepared to argue (module 7), or “we do not know.” Last month is not one of them.
A confounder is a variable that causes both the supposed cause and the supposed effect. It manufactures correlation out of nothing causal, and it does so with total statistical realism: the correlation is genuinely there, reproducible, and significant at any sample size you like.
Meridian’s growth team observes that customers who adopt the mobile app retain at roughly twice the rate of those who do not, and proposes a campaign to drive app installs. Work the mechanism. Engaged customers — the ones who log in, configure things, care about the product — both install the app and stay. Engagement causes adoption; engagement causes retention. The observed adoption–retention correlation flows entirely through those two edges. Drive installs among the disengaged and you get installs, not retention.
The phrase to distrust is “we controlled for that.” Statistical adjustment closes a backdoor path only for confounders you measured, measured well, and thought of in advance. Engagement is a construct; your proxy for it is logins per week, which captures perhaps half of it. The unmeasured half keeps the backdoor open, and no amount of regression output tells you how wide.
A confounder is the alternative explanation the opposing party gets to argue. Your adjustment is your rebuttal to the explanations you anticipated. Randomization is different in kind: it forecloses every alternative explanation at once, including the ones nobody thought to raise — which is why it is a guarantee and adjustment is only an argument.
Confounding is about a third variable. Selection bias is about the door: who ends up in your dataset is itself an outcome, and if the treatment influenced who walked through, the comparison is no longer between comparable groups.
Meridian ran Draft Assist as an opt-in beta first. The 38 agents who volunteered showed CSAT six points above the rest of the floor. The obvious read is that the feature raises satisfaction. The other read is that agents who volunteer for a new AI tool skew senior, skew engaged, and skewed high on CSAT before the beta existed. Seniority causes both opt-in and CSAT; the beta–CSAT association is manufactured by who chose to be measured.
A subtler version conditions on something downstream of the treatment. “Among escalated tickets, drafted replies resolve slower” sounds like an indictment of Draft Assist. But escalation is itself affected by the feature: if drafts resolve the easy cases cleanly, the drafted tickets that still escalate are the genuinely hard residue, while the undrafted escalations include cases that escalated for want of a good first reply. You filtered on an effect of the treatment and compared what was left.
Any comparison where membership in the analyzed group was chosen — by users, by agents, by a filter applied after the treatment — is a selection comparison until proven otherwise. Ask: what process decided who is in this dataset, and could the treatment have touched that process?
This is the sampling objection: the exhibits were selected by the party offering them. You would never accept “here are the twelve emails we chose to produce” as a fair picture of the correspondence. An opt-in beta is the same document production, run by volunteers.
Survivorship bias is selection with a timer. The analysis runs on units that lasted long enough to be measured, and the ones that failed are not merely underrepresented — they are structurally absent, invisible to a query that starts from the current customer table.
Meridian’s customer-success lead observes that the longest-tenured accounts nearly all onboarded with white-glove setup, and proposes white-glove for everyone. The churned white-glove accounts are not in the room. If white-glove was disproportionately given to large, complex accounts, and complex accounts churn harder when they churn at all, the surviving population can look like a white-glove endorsement while the full cohort shows nothing.
Wartime analysts examined returning bombers, mapped the bullet holes, and proposed armoring the areas most hit. Abraham Wald pointed out that the sample consisted entirely of planes that made it back: the unhit regions were where the missing planes had been hit. Armor the engines. The tell is identical to the white-glove claim — the data was collected from survivors, so the pattern of damage is a map of survivable damage.
The operational tell is a query shape: any claim computed from current customers, current employees, retained cohorts, or shipped features carries survivorship risk. The corrective is to define the denominator before you look at the numerator — the full cohort as of its entry date, exits included — and to notice when you cannot, because the exits were never logged.
Module 1 ended with a demand: to license a causal verb you need a credible stand-in for the counterfactual. Randomization is the only method that builds one by construction rather than by argument. This module is about what randomization actually buys, the four decisions that determine whether you collect on it, and the two failure modes — unit mismatch and interference — that void the purchase without producing any visible error.
The through-line is the Draft Assist experiment at Meridian: 240 agents across billing, cancellations, and disputes queues, roughly 12,000 contacts per week. By the end of this module you will have the design one-pager for it, minus the sample size, which module 3 computes with real arithmetic.
Randomization assigns units to arms by a mechanism unrelated to anything about the unit. That single property does the work: because the coin does not know an agent’s tenure, an account’s size, or a contact’s difficulty, no such variable can systematically differ between arms except by chance you can quantify. Every backdoor path from module 1 is severed at once — the measured ones, the unmeasured ones, and the ones nobody has thought of.
Stated formally this is exchangeability: the arms are interchangeable in expectation before treatment, so an outcome difference afterward has exactly one systematic explanation left. Stated practically: randomization is the only technique in this guide whose validity does not depend on your imagination.
Adjustment handles the confounders you named. Randomization handles the confounders you did not name — which, since you cannot enumerate what you failed to think of, is the entire reason it sits at the top of the evidence hierarchy.
Randomization is a mechanism, and mechanisms break. The calibration ritual is the A/A test: run the full pipeline with both arms receiving the identical experience, and confirm the analysis finds nothing at roughly the advertised rate. An A/A that produces significant differences on 3 of 10 metrics is telling you the bucketing is correlated with something — assignment sticky by session, a caching layer serving one arm stale content, a logging join that drops rows unevenly. Better to learn that on a null test than to spend it on a real one.
The unit of randomization is the entity the coin flip lands on. For Draft Assist there are three candidates, and the choice among them determines the experiment’s power, the integrity of its exposure, and its vulnerability to contamination — all three at once, in tension.
Contact-level. Each incoming contact independently gets drafts or not. Maximum statistical resolution: 12,000 independent-ish observations a week. But an agent works both kinds of contact in the same shift, so treatment is not cleanly isolated to the treated units — an agent who learns better phrasing from drafts carries that phrasing into their undrafted replies.
Agent-level. Each agent is drafts-on or drafts-off for the whole run. Exposure is clean and the treatment matches how the feature would actually ship. But outcomes within one agent are correlated — a fast agent is fast on all 300 of their contacts — so 240 agents are worth vastly less than 24,000 contacts, an arithmetic module 3 puts a price on.
Queue-level. Whole queues are assigned. Contamination is nearly eliminated, and so is any hope of precision: three queues means an effective sample of three.
The binding resolution. Draft Assist randomizes at the contact level. Drafts appear only on treated contacts; the agent-learning spillover is documented as a threat, and its direction reasoned out rather than waved at: an agent who improves from exposure to drafts improves on their untreated contacts too, which lifts the control arm and shrinks the measured gap. The bias runs toward the null, so a win survives it and a null result is genuinely ambiguous. That asymmetry is what makes the choice defensible. Module 3 shows the alternative arithmetic: the agent-level design cannot answer the 5% question at any run length, which turns a preference into a constraint.
Symptom: the design doc says “randomized by agent” and the readout says “n = 24,000 contacts, p = 0.01.” Mechanism: contacts within an agent are correlated, so they carry less information than independent observations; treating them as independent shrinks the standard errors. What it fabricates: significance. Not a biased estimate — a confidence interval several times too narrow, and a p-value with no relationship to any error rate. Corrective: analyze at the unit of randomization (agent means), or use standard errors clustered by agent, and say which in the readout.
Every experiment carries an assumption so quiet it usually goes unstated: my treatment does not affect your outcome. Violations are called interference (SUTVA violation), and they are the most common way a technically flawless experiment measures the wrong thing.
At Meridian the mechanism is the shared queue. Treated agents resolve contacts faster, so they return to the queue sooner and pull more work. The queue drains differently. Control agents now face a shorter backlog, a different mix of aged versus fresh contacts, and less time pressure. Their resolution time changes — not because of the feature, but because of other people’s feature. The control arm has stopped estimating the no-feature world.
Reason out the direction rather than asserting it. If control agents face a lighter, fresher backlog, their resolution times improve, the gap narrows, and the experiment understates the feature. If instead the treated agents skim the easy contacts and leave control agents a harder residue, control times worsen and the experiment overstates. Which mechanism dominates is an empirical question you can partly answer — compare control-arm contact difficulty and queue age against the pre-period.
The mitigations, in ascending cost: measure the spillover (track control-arm queue characteristics as a diagnostic), isolate routing so arms draw from separate contact pools, or randomize at the queue or site level and accept the power loss. Meridian takes the first: contact-level randomization means the queue is shared by construction, so the design documents the coupling, monitors the control arm’s workload composition, and treats the null-ward bias as a conservative feature of the win case.
Interference is a side channel. Two processes are supposed to be isolated; they share a resource; state leaks through the resource rather than through the interface anyone is watching. You already know how the analysis goes — enumerate what the two arms share, and treat each shared resource as a candidate channel until you have argued it is not one.
Assignment is not exposure. Of the contacts assigned to Draft Assist, some fraction of agents will never open the drafted reply — busy, skeptical, or working a contact type where the draft is obviously useless. Those contacts were treated by the coin and untreated in fact, and what you do about that determines what question your experiment answers.
Intention-to-treat analyzes every unit by its assignment, regardless of what it actually experienced. This feels wrong to engineers and right to lawyers: it deliberately includes the failures of adoption in the estimate. It is also the only analysis that preserves what randomization bought, and it answers the question you actually face — what happens to resolution time if we launch this feature to the floor, including the agents who ignore it?
The seductive alternative is per-protocol: analyze only the contacts where the draft was opened and used. That comparison is not randomized. Agents choose which drafts to use, and they choose the good ones on the tractable contacts. You have reconstructed, precisely, the self-selection that module 1 spent thirty minutes killing — with the added indignity that the number will look better, which is why someone will ask for it.
If only 70% of assigned contacts are genuinely exposed, the intention-to-treat estimate is roughly 70% of the effect on the exposed. Plan for it: an effect worth 5% among users shows up as 3.5% overall, and an MDE set at 5% will miss it. Either set the MDE against the diluted effect or fix adoption before testing. Discovering dilution after the run is discovering you designed for the wrong question.
Eligibility is the third piece and belongs in writing before launch: which contacts can enter the experiment at all. Meridian excludes contacts routed to the fraud-review queue (different tooling, different SLAs) and contacts opened before the launch timestamp. The rule matters less than its timing — an eligibility filter written after the data exists is a subgroup selection, which module 5 names and prices.
Every discipline in the rest of this guide depends on one artifact existing before the data does. Without it, module 5’s corrections are unenforceable, because there is no record of what was promised. The design one-pager is short by intention — six commitments, one page, circulated and dated.
The last item deserves its own paragraph because it is the cheapest bug-catcher in experimentation. A sample ratio mismatch is an arm-size split too improbable under the intended ratio. On 24,000 contacts split 50/50, the standard deviation of the arm count is about 77, so a split of 12,000/12,000 versus 12,150/11,850 is unremarkable — but 12,552/11,448 is more than seven standard deviations out, a probability with a lot of zeros in it. Nobody gets that unlucky. Something is dropping units non-randomly: a bot filter that fires more in one arm, a logging pipeline that loses treatment rows on timeout, an eligibility check evaluated after assignment. Whatever it is, it is correlated with the treatment, which means the arms are no longer comparable and every downstream number is contaminated. Check it first, read metrics second.
Module 2 left one blank in the design one-pager: how many contacts, for how long. Filling it in is the most consequential arithmetic in experimentation, and it is arithmetic a product manager can do on a whiteboard. This module makes you able to.
The deeper argument here is not about sample size tables. It is that an underpowered experiment is not a weaker version of a good experiment — it is a different instrument with different failure behavior. It misses most real effects, and when it does report a win, that win is systematically too large. A program that runs underpowered tests and ships on their significant results does not merely learn slowly. It learns wrong, confidently, with a chart.
An experiment can fail in exactly two directions. The Type I / Type II error pair names them: a false alarm, where the change does nothing and you conclude it works, and a miss, where the change works and you conclude nothing. Convention fixes the false-alarm rate at α = 0.05 and then, remarkably, leaves the miss rate β to whatever the sample size happens to produce.
That convention is not neutral. Fixing α while letting β float is a rule that protects the status quo: it makes it hard to wrongly adopt a useless change and says nothing about how often you wrongly abandon a good one. In a regulated control environment this asymmetry is often correct. In a product portfolio deciding which of forty ideas to pursue, it quietly sets your miss rate to a number nobody chose.
Statistical power is the complement of the miss rate: the probability that your experiment produces a significant result if the effect is as large as your minimum detectable effect. It is a property of the design — sample size, variance, MDE, α — not of the result, and it is fully knowable before launch. At 30% power, an experiment on a genuinely working feature comes back “no effect” seven times in ten. Run four such experiments on four real wins and you will ship one.
Four quantities are locked in a single relationship: sample size, minimum detectable effect, α, and power. Fix any three and the fourth is determined. There is no fifth lever and no way to buy resolution without paying in one of the others.
For a difference in means between two equal arms, the working formula is:
n per arm = 2 · σ² · (z_α/2 + z_β)² / δ²Every symbol, in words a skeptic could cross-examine: n is the number of units in each arm. σ is the standard deviation of the metric across units — how spread out resolution times are, the noise you are trying to hear through. δ is the minimum detectable effect in the metric’s own units — the smallest true difference you have designed to catch. zα/2 = 1.96 is the strictness of the significance threshold at a two-sided 5%. zβ = 0.84 is the demand that you catch the effect 80% of the time when it is real. The structure says everything: sample scales with the square of the noise and inversely with the square of the effect you want to see. Halving the MDE quadruples the sample.
σ = 340 minutes. δ = 13 minutes (5% of the 260-minute baseline). α = 0.05 two-sided, power = 0.80.
n = 2 × 340² × (1.96 + 0.84)² / 13²
= 2 × 115,600 × 7.84 / 169
≈ 10,700 contacts per armMeridian handles 12,000 contacts a week, split 50/50, so one week yields 6,000 per arm — not enough. Two full weeks yields 12,000 per arm, which clears 10,700 with margin and lands at power 0.84. Verdict: run two full calendar weeks. Whole weeks, because contact mix swings hard by day of week and a run ending on a Wednesday weights the mix.
Now the same question under the design module 2 rejected. Agent-level randomization gives 120 clusters per arm, each contributing roughly 300 contacts over a six-week run. Correlated observations within an agent are discounted by the design effect, 1 + (m − 1) × ICC, where m is contacts per cluster and the intraclass correlation (ICC) — how similar one agent’s contacts are to each other — runs about 0.06 for handle-time metrics.
design effect = 1 + (300 − 1) × 0.06 ≈ 19
effective n = 36,000 / 19 ≈ 1,900 per arm
MDE floor at that n ≈ 31 minutes ≈ 12%Run it twelve weeks instead of six and m doubles, but so does the design effect, and effective n converges on clusters ÷ ICC = 120 ÷ 0.06 ≈ 2,000. The floor barely moves. This is the crucial and counterintuitive result: under cluster randomization, duration buys almost nothing once the clusters bind. The agent-level design cannot answer a 5% question at any run length. Only more agents, a larger MDE, or a less variable metric changes the answer.
Power is not a property of how long you ran. It is a property of how many independent units you observed. When units are clustered, adding time adds observations without adding independence, and the experiment plateaus at a resolution set by the number of clusters and their internal similarity.
The familiar cost of low power is missed wins, and it is the least of the three. At 30% power you miss 70% of your genuine improvements — expensive, but at least the failure is visible as a stack of null results and honestly interpretable as “we could not tell.”
The second cost is not visible at all. Significance is a filter, and at low power the filter only passes estimates that noise has inflated. Consider the underpowered case directly: if your design can only reach the threshold when the measured effect is 6% or larger, and the true effect is 2%, then every result you are permitted to call a win is at least three times the truth. This is exaggeration (Type M) error, and it is not a bias in the estimator — an underpowered experiment is unbiased over all its outcomes. It is a bias in the subset you are allowed to notice.
The third cost is the one that should change how you read your own program’s track record. Significance answers “how surprising is this data if nothing is happening,” but the question you care about is “given that this came back significant, how likely is it real?” — and that depends on how many of the ideas you test are real in the first place. Suppose 10% of the changes your team ships are genuine improvements, an honest and even generous base rate.
At 80% power: out of 100 tested ideas, 10 are real and you catch 8; 90 are null and you falsely flag 4.5. You report 12.5 wins, of which 4.5 — 36% — are noise. At 20% power: you catch 2 of the 10 real ones, still falsely flag 4.5 of the nulls, and report 6.5 wins of which 4.5 are noise — 69%. Most of that program’s celebrated wins never happened.
The argument you will actually face is not that power does not matter. It is: “We know it is small, but some evidence beats none — let us just run it and see.” The arithmetic above answers it. An underpowered test consumes traffic and calendar, produces a number that carries the visual authority of evidence, and then contaminates the decision record with a result that is either a miss you will mistake for a refutation or a win you will mistake for a large effect. No evidence at least leaves the question open and the team honest about the state of its knowledge.
There are four legitimate responses when the calculation says you cannot reach power, and each is a real option rather than a consolation:
One cheap design is legitimate: a pre-declared directional gate — “we will roll back if the point estimate is negative, accepting that this decision rule has roughly a one-in-three error rate at this sample size.” It is legitimate precisely because the error rate is computed and printed in the readout before the run. The same rule invented afterward is the anti-pattern module 4 dismantles.
Symptom: a design doc with no sample-size calculation, or a run length set by the sprint calendar rather than by δ and σ. Corrective: make the power calculation a required field in the one-pager, and when it fails, force the choice among the four options above in writing. The failure mode is not the missing math; it is that nobody had to say out loud which compromise they were making.
The experiment has run. What arrives is a number, an interval, a p-value, and a slide. This module is about reading them without importing the folklore that comes attached — folklore that survives in otherwise sophisticated organizations because the correct statements are one conditional-probability step away from the incorrect ones, and nobody is checking which way the arrow points.
You have a professional advantage here. The most common misreading of a p-value is a fallacy you have already been trained to catch under a different name, in a room with higher stakes. This module leans on that training hard.
Here is the sentence, and it repays memorizing: a p-value is the probability of observing data at least this extreme, if the null hypothesis were true. It is a statement about data under an assumption. It is the probability of the evidence given no effect. It is never the probability of no effect given the evidence.
You already own the correction. In a criminal case, an expert testifies that the probability of a random innocent person matching the recovered DNA profile is one in a million. Opposing counsel restates this as a one-in-a-million chance the defendant is innocent. That restatement is the prosecutor’s fallacy, and it is wrong for a reason that has nothing to do with DNA: it inverts the conditional. Converting P(match | innocent) into P(innocent | match) requires a further ingredient the testimony never supplied — how many people were in the candidate pool to begin with. In a city of ten million, roughly ten innocent people match.
The folklore, dismantled item by item. “p = 0.03 means a 3% chance the result is due to chance.” No: it means that if the result were due to chance alone, data this extreme would appear 3% of the time. “p = 0.049 is a win and p = 0.051 is a loss.” No: these are the same evidence, and the threshold is an administrative convention that turns a continuous measure into a binary for the convenience of decision-making. “a very small p means a big effect.” No: p depends on effect size and sample size, so at two million observations a 0.02% change yields a spectacular p-value and no reason to do anything. “p > 0.05 means the feature does nothing.” No, and this one is the most expensive; it is the subject of the next section.
A confidence interval is the set of true effect sizes the data cannot rule out. That formulation is doing real work: it turns the result from a verdict into a bounded claim about what remains possible, which is exactly the form a decision needs.
Width is information, and it is the information a p-value discards. Two results can both be “not significant” and mean opposite things. An interval of [−15.9%, +3.5%] says the data is compatible with a large improvement and a modest regression — you know almost nothing. An interval of [−1.5%, +0.7%] says the true effect is within a point of zero in either direction — you know a great deal, namely that this change does not move this metric. The first is ignorance. The second is a finding. Only the second licenses the sentence “this change does nothing,” and no p-value can distinguish them because both report p far above 0.05.
Absence of evidence is not evidence of absence — unless the interval is tight. A narrow interval around zero is how you convert a non-significant result into a real finding, and it is the only honest route to the sentence “this does nothing.”
Significance is about surprise. Decisions are about magnitude. A result can be overwhelmingly significant and completely irrelevant: at two million contacts, a 0.02% change in CSAT will produce p = 0.001 and justify precisely nothing, because 0.02% of anything is not worth an operational change. The reverse also happens — a 9% improvement with p = 0.11 at small n is a serious candidate for a larger test, not a rejected hypothesis.
Read the two Draft Assist scenarios as the discipline in action.
Two full weeks, contact-level, 12,000 per arm. Mean resolution time falls 21 minutes, −8.1%, 95% CI [−11.4%, −4.8%], p < 0.0001. Guardrails: CSAT flat with an interval excluding any meaningful decline, reopen rate flat, compliance-flag rate flat. Reading: the interval excludes zero and excludes everything weaker than −4.8%, which is comfortably above the 5% MDE, so the effect is not merely real but decision-grade at its lower bound. The guardrail intervals are tight enough to exclude harm rather than merely failing to detect it — a distinction the previous section makes load-bearing. Decision: ship, and forecast on the lower bound, not the point estimate.
A follow-up variant, one queue, one week: 1,400 contacts per arm. Resolution time −16.2 minutes, −6.2%, 95% CI [−15.9%, +3.5%], p = 0.21. The deck reads: “directionally positive, recommend ship.” Reading: the data cannot distinguish a 16% improvement from a 3.5% regression. The point estimate is the least informative number on the slide, and — per module 3 — if it had cleared significance at this sample it would have been inflated. Decision: neither ship nor kill. Extend to a powered sample, or redesign the variant. Record that the experiment answered nothing, which is itself worth knowing about the design.
Symptom: a readout leads with the sign of a point estimate and omits the interval; the recommendation is “trending well, let us ship and monitor.” Why it persists: it feels like appropriate pragmatism, and it is unfalsifiable — every experiment’s point estimate has a direction. Corrective: the decision standard must be a pre-set interval criterion — “ship if the upper bound is below −2%” — declared in the one-pager. Then “directionally positive” either meets the criterion or does not, and the phrase stops doing work it was never entitled to do.
This is usually presented as a religious war. It is better understood as two instruments answering two different questions, each with a real cost.
The frequentist offer. Guaranteed long-run error rates: design the procedure and, whatever the truth, you will falsely adopt at most 5% of null changes and detect real ones at your stated power. Nothing needs to be assumed about how likely the effect was beforehand, so nothing about your beliefs is available to be litigated by a stakeholder with an agenda. That property is exactly what a ship/no-ship gate and an audit record want. The price: it answers an oblique question. It tells you how surprising the data would be under a hypothesis nobody believes literally, and it will not tell you the probability that the feature works, because in this framework the effect is a fixed unknown constant and does not have a probability.
The Bayesian offer. A direct answer to the decision-maker’s actual question: given a stated prior and this data, the probability that the effect exceeds the threshold you care about — for instance, an 82% probability the improvement beats 5%. That number composes correctly into an expected-value calculation and updates coherently as data arrives, which makes it the natural instrument for portfolio judgment and continuous monitoring. The price: a prior you must state and defend. This is also its honesty. Everyone brings a prior to a results meeting; the Bayesian writes it down where you can attack it, while the frequentist smuggles it in through the phrase “that seems too good to be true.”
Working guidance. Use the frequentist gate for ship/no-ship decisions and the compliance record, because guaranteed error rates are what an audit needs and priors are what an audit fights about. Use Bayesian reasoning for portfolio allocation, for continuous harm monitoring, and for any question phrased as “how confident should I be.” Where they disagree materially, the disagreement is a signal that the data is weak and the answer is more data, not a better framework. And distrust anyone who switches frameworks after seeing the result — that is not statistics, it is shopping.
The frequentist gate is a burden of proof set before trial: a fixed standard, applied identically regardless of who the defendant is. Bayesian reasoning is how a judge actually updates through a hearing. Both are legitimate; the misconduct is choosing which one governs after the evidence is in.
Everything so far assumed good faith and competence. This module assumes both and shows that they are not enough, because each deception in the catalog is a locally reasonable act: checking on your experiment, looking at more than one metric, noticing a pattern in the data. The damage comes from what those acts do to the error rate you thought you had bought — and the error rate does not announce its own inflation. The readout looks identical.
Four entries, one root. Each one moves the goalposts after the ball is in the air: the stopping rule, the metric, the hypothesis. Each has a symptom you can spot in a slide, arithmetic that shows the cost, and a correction that is cheap if committed to in advance and impossible afterward.
The experiment is live. The dashboard updates hourly. Someone checks it every morning, and on the day it crosses p < 0.05, the team ships. Every step of this is natural, and the combination destroys the guarantee the 5% threshold was supposed to provide.
The mechanism is easiest to see if you picture the p-value over time. It does not descend smoothly toward the truth; it wanders, because each day’s data adds noise as well as signal. Under a genuinely null experiment, the trajectory random-walks, and a random walk watched at fourteen points has fourteen chances to dip below any line you draw. The 5% guarantee was priced for one look at a pre-specified time. Buy fourteen looks and pay fourteen times — the real false-alarm rate for daily checks over a two-week run lands somewhere around 25%, roughly five times the advertised rate.
The correction costs nothing but discipline, and module 2 already paid for it: the horizon was fixed in the design one-pager before launch. Two additional practices make it enforceable. Report the look count in every readout — “this analysis is the single pre-registered look at day 14” is a sentence that either can or cannot be written. And treat extending a run because the result is not yet significant as the same offense wearing a different hat: it is a stopping rule that depends on the data, which is the definition of the problem.
Symptom: the ship decision date is earlier than the horizon in the design doc, and the p-value is just under the line. Mechanism: the threshold was priced for one look; the team took seven, and stopped at the most favorable one. Corrective: read at the horizon, or use a sequential design that has budgeted for the looks in advance. Cost of the corrective: waiting, and occasionally watching a real win sit unshipped for a week.
The prohibition on peeking is unsatisfying, because looking early is genuinely valuable — you want to catch harm, and you want to stop a clear winner sooner. The resolution is that you may peek if you pay, and the payment is arranged in advance.
Group-sequential designs pre-specify the looks — say, at days 4, 7, 10, and 14 — and assign each a stricter threshold, so that the total false-alarm probability across all four still sums to 5%. This is alpha spending: a fixed budget of surprise, allocated across looks rather than spent at one. Early looks demand dramatic evidence (p below roughly 0.001), which is exactly right — stopping after three days should require a result so strong that noise is an implausible explanation. The cost is real: if you spend budget early and do not stop, the final look is slightly stricter than a single-look design would have been, which costs a little power.
Always-valid inference is the stronger version — confidence sequences that remain correct no matter how often you look — and it is worth naming because it is a platform capability to ask your experimentation team for rather than something to implement yourself. If your platform offers it, continuous monitoring becomes safe by construction and this entire section becomes someone else’s problem.
Peeking rules protect the ship decision. They do not protect the customer. Monitoring a guardrail for harm and stopping early when it breaches is legitimate, pre-planned, and mandatory — the error you are guarding against there is failing to stop, not stopping too eagerly. Meridian’s compliance-flag guardrail is monitored sequentially with a documented stopping rule from day one, and nobody invokes “no peeking” to keep a harmful variant live.
Run one test at α = 0.05 on a change that does nothing, and you have a 5% chance of a false winner. Run twenty independent tests on the same nothing and the chance that at least one comes back significant is 1 − 0.95²⁰ ≈ 64%. The victory-lap deck — “resolution time was flat, but logins are up 3% and NPS moved two points!” — is not evidence of a hidden benefit. It is the expected output of a null experiment with a wide dashboard.
The corrections, in order of usefulness. First and most important: one pre-registered primary metric. The other nineteen are context, diagnostics, and guardrails — they inform your understanding and can veto, but they cannot promote a null result into a win. This single rule handles most of the damage, and it is free.
Second, for the cases where you genuinely have several co-equal hypotheses, apply a correction to the family: Bonferroni divides the threshold by the number of tests (simple, strict, costs power), or false-discovery-rate control targets the proportion of your declared winners that are false rather than the chance of any error at all (more forgiving, better suited to screening many candidates). The choice is a policy question: are you trying to avoid any false claim, or to keep the false share of your claims tolerable?
Subgroup fishing is the same sin with demographics instead of metrics. “It works for enterprise accounts in EMEA on mobile” multiplies the tests by the number of slices you were willing to cut, and nobody ever counts the slices they looked at and abandoned. If subgroups matter to the decision, name them in the one-pager and treat them as a pre-registered family with a corrected threshold.
The fourth entry is the subtlest, because it involves no statistical error at any single step. HARKing — Hypothesizing After Results are Known — is constructing the hypothesis from the data and then presenting that same data as its confirmation. Every individual move is legal. The composition is circular.
It shows up in a specific narrative shape. The experiment comes back flat. Someone digs, finds that the effect appears concentrated among new customers, and reconstructs a plausible story: new customers have no established workflow, so they adopt the drafted reply rather than fighting it. The story is good. It may even be true. What it cannot be is tested by the data that generated it, because the data was searched over many possible stories and this one was selected for fitting.
A theory built to fit the evidence cannot then cite that evidence as independent corroboration. You would take that objection apart in cross without preparation: counsel constructed the timeline after reviewing the documents, then offered the documents as proof of the timeline. The statistical version has better graphics and the same defect.
The honest disposal route is not to suppress the finding — exploratory analysis is valuable and often where the good hypotheses come from. It is to label it correctly and route it to the next experiment. Say: “Exploratory: the effect appears concentrated among accounts under 90 days old. This was not pre-registered and the subgroup was one of eleven examined. We propose a confirmatory test on new accounts with an MDE of 8%, powered at two weeks.” That sentence loses nothing and claims nothing it has not earned.
The instrument that makes all of this enforceable is pre-registration — module 2’s design one-pager, dated and circulated before data exists. Its power is not moral. It is that it makes the difference between a prediction and a postdiction visible, and the deceptions in this catalog all depend on that difference being invisible.
Peeking moves the stopping rule. Multiple comparisons move the metric. Metric fishing moves the outcome definition. HARKing moves the hypothesis. All four are the same act — changing the terms of the test after seeing the result — and all four are prevented by the same artifact, which is a dated document nobody can edit afterward.
Module 5 established that the primary metric is the antidote to most self-deception. This module is about how to choose it, what the other numbers are for, and why a guardrail that can be argued away is not a guardrail. The framing is architectural on purpose: a dashboard is not a list of things worth knowing, it is a decision system with roles, authority, and failure modes.
Two of those failure modes deserve their own vocabulary. Metrics get gamed by the people whose work they measure, reliably and quickly. And metrics lie early, in two directions, for reasons that have nothing to do with your feature working.
Four jobs, and every number on your dashboard holds exactly one of them.
The primary metric is the pre-registered decider. There is exactly one, it is chosen before the run, and the ship decision rides on it alone. Its defining test: would we ship on this metric alone if everything else were unchanged? If the answer is no, it is not your primary.
A guardrail metric is defended, never optimized. It holds veto authority and cannot make the case to ship. Its test: would we block on this alone? A metric nobody would block on is not a guardrail — it is decoration with a serious-sounding name.
A driver metric is a leading indicator of the mechanism. It explains why the primary moved, moves faster than the outcome, and is the first place you look when a result surprises you. Its test: does this explain, rather than decide?
Vanity metrics move with everything and decide nothing — total drafts generated, tokens served, dashboard sessions. They are not evil, they are simply not load-bearing, and the failure mode is that they occupy the top of the deck where a decision-grade number should be.
The guardrail’s authority is deliberately asymmetric: it can stop a launch, and it can never authorize one. That asymmetry is the whole design. A metric with symmetric authority is a second primary, and two primaries is one more than the number of things a single decision can ride on.
Draft Assist prompt v2 is tuned for a warmer, more confident tone. The readout: mean resolution time down 9.0%, interval clearly favorable, well past the MDE. Guardrails: compliance-flag rate 0.8% → 1.9% (+1.1pp, 95% CI [+0.7, +1.5pp]); reopen rate 4.1% → 5.3% (+1.2pp, 95% CI [+0.5, +1.9pp]); CSAT flat.
Mechanism diagnosis. These are not two independent problems and a win. They are one behavior seen three ways. The v2 drafts are faster because they are confidently wrong: they commit to a refund timeline or a policy reading without hedging, agents send them with less editing, the contact closes quickly — and then QA flags the misstatement and the customer writes back. Speed, flags, and reopens are the same phenomenon measured at three points in time.
Decision: veto. A 1.1-point rise in replies that misstate refund terms is a control failure with regulatory exposure; there is no resolution-time number that buys it. The information dividend: the breach identified the mechanism precisely, which tells you what v3 must do — gate drafts behind a compliance classifier that blocks or flags any draft asserting a refund amount, timeline, or policy exception, then re-test. The failed experiment paid for the next design.
Two disciplines make a veto real. First, guardrail thresholds are set before the run, in the one-pager, as pre-specified breach conditions — “any increase in compliance-flag rate whose interval excludes zero halts rollout.” A threshold negotiated after the data arrives is a negotiation, not a control. Second, guardrail intervals must be read for whether they exclude harm, not for whether they reached significance: a flat point estimate with an interval spanning [−2pp, +3pp] on compliance flags has not cleared anything, it has failed to look.
Guardrails are compensating controls, and a breached guardrail is a control finding, not a footnote. You already know what happens to a program that treats control findings as inputs to a cost-benefit conversation with the business owner who wants to ship — the control stops being a control the first time it is overruled, and everyone learns that in one meeting.
You randomize on what moves in weeks; you care about what moves in quarters. The bridge between them is a chain of proxies, and its links are assumptions of wildly different quality that a dashboard renders as identical rectangles.
The characteristic failure is promoting a driver to primary because it is convenient. Draft-open rate moves within hours, has tiny variance, and produces beautiful charts — and measuring it tells you that agents open drafts, not that anything got better. Optimizing a driver directly is how you end up with a feature that maximizes its own usage. The primary must be the nearest metric downstream that would justify shipping on its own.
Goodhart’s law: when a measure becomes a target, it ceases to be a good measure. The mechanism is not mysterious. People whose compensation, ranking, or standing depends on a number will move that number, and moving the number is almost always cheaper than moving the thing the number was a proxy for.
Meridian has already run this experiment without meaning to. When handle time became a floor-level target, it improved 15% in a quarter — and reopen rate climbed, because the fastest way to close a ticket is to close it before the customer is finished. The handle-time improvement was real and the underlying performance was worse.
The design pattern that survives this is metric–guardrail pairs: for every metric you make a target, name the cheapest way to game it and instrument the tripwire that gaming trips. Speed pairs with reopen rate and CSAT. Volume pairs with quality sampling. Automation rate pairs with escalation rate. Cost per contact pairs with repeat-contact rate. The pair is not a nicety; a target without its tripwire is an instruction to game.
Red-team the metric before it ships. Take the adversarial posture you would take toward an authorization model and point it at the incentive: who is measured on this, what is the cheapest action that moves it, does that action serve the customer, and what would detect it? Treat “nobody would do that” as the answer that has never once been correct about an incentive.
Two transients distort early readings in opposite directions, and both resolve toward the truth if you let the experiment run.
Novelty: people engage with a new thing because it is new. Usage spikes, the metric it drives spikes, and the effect decays toward truth from above. Primacy: a change disrupts practiced workflows, so performance dips while people relearn, and the effect recovers toward truth from below. Draft Assist is a primacy-shaped feature: agents who have written replies from scratch for four years must learn a different motion — read, judge, edit — and they are slower at it before they are faster.
The detection method is a shape, not a test: plot the treatment effect by week of exposure rather than calendar week, so that users entering later are aligned by their own experience rather than by the date. If the first week and the last week disagree materially, the run was not long enough, and the honest readout says which direction the effect is still moving.
Symptom: a decision made from the first days of exposure, usually justified by early significance. Corrective: set run length from the transient as well as from power — long enough for the learning curve to flatten — and report the week-one versus week-final comparison in the readout as a standard field. Note the asymmetry of cost: shipping on novelty wastes engineering on an effect that will evaporate; killing on primacy discards a feature that works, and you will never learn that you did.
Randomization is not always available, and the reasons are usually good ones: a regulator or works council forbids differential treatment, a fairness constraint makes withholding a benefit indefensible, a platform change is physically all-or-nothing, or the contamination between arms is so severe that a randomized comparison would measure the mixture rather than the treatment.
What remains is a family of designs that reconstruct the counterfactual from structure rather than from a coin. They are not consolation prizes — a well-argued difference-in-differences beats a badly contaminated experiment — but they shift the burden. In a randomized trial, the identifying assumption is enforced by the mechanism. In a quasi-experiment, it must be argued to a skeptic, and that argument, not the point estimate, is the deliverable.
When the coin flip is off the table, the options are not equally good, and the ordering principle is simple: how much of the identification is guaranteed by design versus supplied by your argument.
| Design | Where the counterfactual comes from | What you must argue |
|---|---|---|
| Randomized experiment | The coin flip | Only that the mechanism worked (SRM, A/A) |
| Natural experiment | An external shock that assigned treatment for reasons unrelated to the outcome | That the shock was genuinely unrelated |
| Designed quasi-experiment (DiD, RD) | Structure: a shared time trend, or an arbitrary threshold | A specific, stateable assumption — parallel trends, no manipulation |
| Adjusted observational | Statistical adjustment on measured covariates | That you measured every confounder that matters |
| Raw correlation | Nowhere | Everything |
The fourth row is where most business analysis lives and why most business analysis is weak: “we measured every confounder that matters” is an unfalsifiable claim about your own imagination, and module 1 showed why more control variables do not strengthen it. The third row is different in kind. Parallel trends and no-manipulation are specific claims with observable implications, which means they can be tested against evidence, attacked by a skeptic, and defended or abandoned on the merits.
This is a burden-of-proof shift, and you should treat it procedurally rather than as a matter of degree. Under randomization, the mechanism carries the burden. Under a quasi-experiment, you carry it — which means the deliverable is not a number but a memo: here is the comparison, here is the assumption it rests on, here is the evidence for the assumption, and here is the single fact that would defeat it.
Meridian cannot randomize Draft Assist across its whole support organization, because its EU works council must approve any tool that monitors or augments agent work, and that approval will not arrive before Q3. Legal forces a staggered rollout: North America gets the feature in March, the EU does not. This is a constraint, and it is also a design.
The naive read is the before/after on North America: 262 → 241, so the feature saved 21 minutes. Module 1 already refused that, because it substitutes February for the counterfactual. But now you have something February never gave you — a group living through the same calendar without the treatment. The EU went 259 → 254, a 5-minute improvement with no feature at all: tooling upgrades, seasonal mix, a maturing support playbook, the ordinary secular drift that would have lifted North America too.
NA: 262 → 241 change = −21 min
EU: 259 → 254 change = −5 min
difference-in-differences = (−21) − (−5) = −16 min (−6.1%)The second difference is the whole trick: it subtracts whatever both regions experienced in common. Whatever the shared drift was — seasonality, the new knowledge base, a quieter billing cycle — it hit both and cancels.
Which brings us to the assumption, stated so a skeptic can swing at it: absent Draft Assist, the gap between North American and EU resolution time would have stayed roughly constant. That is parallel trends. It is not the claim that the regions are similar — they are not; the EU runs three minutes faster at baseline and that is fine. It is the claim that whatever makes them differ is stable across the period.
You argue it with evidence. Meridian has eight pre-period weeks in which the NA–EU gap held within ±3 minutes against a 16-minute estimated effect — a gap that stable does not suddenly move five times its historical variation by coincidence. You also argue it by exclusion: what could break it? Any change affecting one region and not the other during the window. An EU-only migration to a new ticketing system. A North American staffing surge. A regional product launch that changed contact mix. A holiday calendar that falls differently. Each of these is checkable, and the memo should say which ones you checked.
Not a violation of similarity — a violation of parallelism. And note what does not repair it: adding control variables after the fact. If the EU migrated ticketing systems mid-period, no covariate adjustment recovers the counterfactual, because the comparison group stopped being a valid stand-in for the treated group’s trend. The honest response is to shorten the window to exclude the shock, find a cleaner comparison region, or abandon the claim.
Your organization already enforces arbitrary thresholds — spend tiers, tenure cutoffs, score bands, SLA classes — and every one of them is a small experiment nobody ran on purpose.
Meridian assigns a dedicated customer success manager to accounts above $10,000 in annual spend. The design intuition: an account at $9,950 and an account at $10,050 are, in every respect that matters, the same kind of account. Nothing about the business changes across a hundred dollars of spend — except the rule, which fires on one and not the other. In a narrow window around the cutoff, CSM assignment is as good as randomized, so comparing retention just above the line against just below estimates the CSM effect for accounts near the threshold.
Two assumptions, both stateable in a sentence. No precise manipulation: nobody can choose which side of the line they land on. This fails the moment a sales rep learns that nudging an account to $10,001 wins it CSM coverage, and it fails visibly — the density of accounts spikes just above the threshold, which is why the histogram check is mandatory rather than optional. Nothing else changes at the same threshold: if $10,000 also triggers a different support tier, priority routing, and a quarterly business review, the jump measures the whole bundle, and no amount of data separates the CSM from its co-travelers.
The scope limit is important and often forgotten: regression discontinuity estimates the effect near the cutoff. It tells you what a CSM does for a $10,000 account. It says nothing about what a CSM would do for a $200,000 account, where the relationship may be entirely different — and that extrapolation is the most common way a valid estimate gets misused.
The final defense is linguistic, and it is the one you exercise most often, because most claims that cross your desk will never be upgraded to any design at all. Each design licenses a specific sentence, and the discipline is refusing to speak above your evidence.
| Evidence you have | The sentence you may say | The sentence you may not |
|---|---|---|
| Raw correlation | “X and Y move together in our data.” | Anything with a causal verb |
| Correlation + adjustment | “The association persists after adjusting for size, tenure, and queue mix.” | “Controlling for confounders, X causes Y” |
| Difference-in-differences | “Relative to a comparison region on a parallel pre-trend, X was followed by a 16-minute improvement; the estimate holds if no region-specific shock occurred.” | “X caused a 16-minute improvement” without the conditional |
| Regression discontinuity | “For accounts near the $10,000 threshold, the rule raises retention by N points.” | Any claim about accounts far from the cutoff |
| Randomized experiment | “X causes a 21-minute reduction, 95% CI [.., ..], in this population over this window.” | Claims outside the tested population or window |
Two habits round this out. The first is a sensitivity question you can ask without any math: how large would an unmeasured confounder have to be to erase this estimate? If a modest, entirely plausible difference in customer engagement would account for the whole gap, the claim is fragile regardless of its p-value. If it would take a confounder larger than any variable you have ever measured, the claim is robust — and that is an argument a skeptic can engage with rather than a defensive assertion that you controlled for things.
The second is knowing when observational evidence is genuinely enough. It often is: for decisions that are cheap and reversible, for effects far too large for any plausible confounder to manufacture, and for decisions you would make anyway where the analysis only calibrates expectations. The cost of over-demanding rigor is real — a team that will not act without an RCT will be beaten by one that acts well on good observational evidence. The discipline is not refusing to decide. It is refusing to misdescribe what the decision rests on.
LLM features attract two opposite errors. The first is that they are so novel that ordinary experimental discipline does not apply — usually expressed as shipping on eval scores and vibes. The second is that they are identical to any other feature and need no special handling, which walks straight into version drift and per-response variance.
Both are wrong in the same way: what changes is the variance structure and the confounders, not the logic. The counterfactual argument, the power arithmetic, and the entire self-deception catalog apply unmodified. This module covers the three things that genuinely differ — the unit calculus, the offline-to-online linkage, and the version confounders — using Draft Assist’s prompt-v2 test as the worked case.
Four things change, and each has a design consequence rather than merely being interesting.
Nondeterminism. The same contact submitted twice yields different drafts. This is not a bug to be eliminated — temperature is often load-bearing for quality — but it means a single response is a draw from a distribution, and per-response quality metrics carry more variance than a deterministic feature’s would. It also means an A/A test on an LLM feature is genuinely informative: two arms running the identical prompt will differ, and knowing how much they differ is your noise floor.
Version drift. The provider ships a model update mid-run. A teammate improves the system prompt on Tuesday because it obviously needed it. The retrieval corpus gets re-indexed. Each of these changes the treatment while the experiment is measuring it, and none of them appears in your experiment platform. This is the single most common way an LLM experiment silently becomes uninterpretable.
Per-response variance dwarfs per-user variance. Draft quality varies enormously between contacts — a password reset drafts perfectly, a multi-invoice billing dispute does not — and that spread is typically much larger than the spread between agents. That is good news for power on response-level metrics and irrelevant to power on agent-level outcomes, which is a distinction the next section makes operational.
Cost is a first-class metric. Tokens are cost of goods sold, and a variant that improves resolution time by 6% while tripling inference cost may be worse than the control. Cost per contact belongs on the readout as a standing guardrail, not as a footnote after launch.
The counterfactual logic of module 1, the unit and interference reasoning of module 2, the power arithmetic of module 3, the interval discipline of module 4, the entire self-deception catalog of module 5, and the metric architecture of module 6 all apply without modification. LLM features get no exemption; they get extra confounders.
Module 2’s rule holds: match the unit of randomization to the behavioral footprint of the treatment. LLM features make this sharper, because the same product can host changes with radically different footprints.
Response-level. Each individual draft is independently generated by prompt v1 or v2. Correct when the change is invisible to workflow — prompt wording, model swap, retrieval tweak, temperature. The agent cannot tell which variant produced the draft in front of them, so there is nothing to learn and nothing to contaminate. Maximum power: Meridian generates 12,000 drafts a week, and per-response metrics like edit distance and draft acceptance stabilize within days.
Contact- or session-level. Correct when the treatment shapes a whole interaction — a multi-turn assistant, a change that affects the second reply given the first. Randomize the conversation, not the turn, or the arms interleave inside a single customer experience.
Agent-level. Required when the treatment changes workflow or trains behavior. The original Draft Assist launch was agent-visible: the feature’s existence changes how an agent works, and an agent who has drafts on half their contacts learns differently than one who has them always or never. Its prompt version, by contrast, is invisible — which is why the v2 test runs response-level on a fraction of traffic while the launch test could not.
The trap in the dashed band deserves naming because it is easy to walk into while optimizing for power. If you randomize drafts on and off at the response level and measure resolution time, every agent experiences a 50/50 mixture. They adapt to the mixture — hovering, waiting to see whether a draft appears, developing a hybrid workflow. Neither arm reflects the world where the feature is always on or always off, and the measured difference is a within-mixture contrast that will not survive launch.
An offline eval is a versioned test suite for model behavior: fixed inputs, scored outputs, run in CI before anything reaches a customer. Guide Nº 03 covers how to build one. The question here is what an eval score means in the evidence hierarchy this guide has built.
The answer is precise: an eval is a driver metric. It sits at the far left of module 6’s proxy chain, it moves in minutes rather than weeks, it explains rather than decides, and — like every driver — it earns its position only by predicting the outcome it proxies for. The division of labor is the familiar one from software: evals are the unit tests, cheap and fast and run constantly; the online experiment is the integration test against reality, expensive and slow and authoritative.
The linkage discipline is concrete. Across your last several launches, plot the eval delta against the measured online effect. If the relationship is real, the eval has earned its gate authority and you can screen aggressively with it. If the scatter is a cloud — and it often is, for evals assembled from cases that were easy to write rather than cases that predict outcomes — then the eval is measuring what someone thought mattered, and it needs rebuilding against cases drawn from real traffic.
When they disagree on a specific launch, the resolution is not symmetric. Reality wins: the experiment measured what customers and agents actually did, and the eval measured a proxy. The eval gets recalibrated — add the cases the online result revealed, re-weight the scoring — and it keeps its job as the cheap pre-filter. What does not happen is the experiment being re-run until it agrees with the eval.
Symptom: a launch recommendation whose evidence section contains only offline scores, usually for a change described as “just a prompt tweak.” Corrective: require an online result for anything that reaches a customer, and set the bar by blast radius — a response-level prompt test costs days, which is affordable for nearly every change worth making. Where the change is genuinely too small to test, say so explicitly and record it as an untested ship.
Two practical matters close the module: how long these tests need to run, and what must be frozen while they do.
Duration follows the metric, not the feature. Per-response quality metrics — draft acceptance rate, edit distance between the draft and the sent reply, hallucination-report rate — have large per-response variance but enormous sample counts, and they stabilize in days. Agent-behavior and customer metrics — resolution time, CSAT, reopens — obey module 3’s arithmetic unchanged, which for Meridian means two full weeks at contact level. A prompt-v2 test can therefore report acceptance and edit distance after three days and must still wait for resolution time, and a readout that presents the fast metrics as if they settled the slow question is the driver-as-primary error in a new costume.
Freeze the treatment. Three things must be pinned for the run’s duration and stated in the one-pager: the model version (pin an explicit snapshot identifier, not a floating alias; a provider upgrade mid-run is an uncontrolled treatment applied to both arms at once), the prompt and its template (no improvements during the run, however obvious — an edit on day 4 means days 1–3 and 5–14 tested different things), and the retrieval corpus version if the feature retrieves. Log the version identifiers alongside every response so that if something moves anyway, you can segment the analysis at the boundary instead of discarding the run.
Guardrails for a generation feature extend module 6’s set with the failure modes specific to generated text: compliance-flag rate (already Meridian’s sharpest instrument, since it detects exactly the confident-and-wrong failure), hallucination-report rate from agents, escalation rate, and cost per contact in tokens. Each is defended, none is optimized, and each has a pre-specified breach condition.
Guide Nº 03 builds the eval suite this module points online. The relationship runs both directions: an eval that has never been checked against an online result is an untested proxy, and an experiment program with no eval layer pays full experimental cost for changes that a three-minute offline run would have rejected.
Eight modules of machinery reduce to one capability: reading a results readout and knowing, within a few minutes, whether it can carry the decision someone is about to make with it. This module is the reference layer — the catalog in one table, the argument that null results are assets, and the ten questions that turn a slide into testimony.
None of this requires you to run the analysis. It requires you to know which questions are lethal and why, which is the skill your first career was built on and the reason this guide has leaned on it throughout.
Six patterns account for most of the bad decisions made from good data. Each is stated as a symptom you can spot without access to the underlying analysis, because that is the situation you will usually be in.
| Anti-pattern | Symptom in the wild | Corrective | Taught in |
|---|---|---|---|
| Shipping on a peeked p = 0.049 | The decision date precedes the documented horizon; the p-value sits just under the line | Read at the horizon, or specify a group-sequential boundary before launch; report the look count | Module 5 |
| “Directionally positive” as a decision standard | The recommendation cites the sign of a point estimate; no interval appears on the slide | A pre-registered interval criterion — ship only if the upper bound clears a stated threshold | Module 4 |
| Twenty metrics and a victory lap on the two that moved | The primary is flat; secondary metrics carry the recommendation | One pre-registered primary decides; corrected thresholds for genuine families; count the subgroups examined | Module 5 |
| Randomizing agents, analyzing contacts | The design says one unit and the sample size says another, orders of magnitude apart | Analyze at the randomization unit or cluster the standard errors, and say which | Module 2 |
| Ignoring interference on a shared queue | The control arm moves during the run; nobody asks why | Enumerate shared resources; monitor control-arm conditions; isolate routing or reason out the bias direction | Module 2 |
| Treating a failed experiment as a failed feature | A null result ends a program; nobody states what the interval ruled out | Read the width: a tight null is a finding, a wide null is a request for more information | This module |
Notice the shape of the correctives. Five of the six are artifacts — a document, a threshold, a field in a template — rather than acts of vigilance. That is deliberate. Vigilance does not survive a quarter under deadline pressure; a required field in a readout template does.
A null result with a tight interval is a finding, and treating it as a failure is the most expensive habit in this guide because it corrupts what a team is willing to test.
Draft Assist v3 returns −0.4%, 95% CI [−1.5%, +0.7%]. That is not “we learned nothing.” It is: this variant does not move resolution time by more than 1.5% in either direction, and we will not be revisiting that hypothesis. A roadmap myth has been killed cheaply and permanently. Compare the alternative — shipping it on a wide interval and spending two quarters attributing unrelated fluctuations to it.
The portfolio frame follows directly from module 3’s arithmetic. If roughly one idea in ten genuinely works, a well-run program has a low win rate by construction. A team reporting that eight of nine experiments were wins is not outperforming the base rate by a factor of eight; it is exhibiting the catalog — peeking, metric flexibility, post-hoc subgroups — or testing changes so certain they did not need an experiment. Either way the number is a symptom, and the correct response to it is curiosity about their protocol, not envy.
The experiment refuted this implementation’s effect on this metric, in this population, over this window. It did not refute the hypothesis, the mechanism, or the next variant. Popper’s asymmetry cuts both ways: a refutation is decisive about exactly what was tested and silent about everything else. Teams that forget the first half ship on noise; teams that forget the second half kill good ideas on one bad implementation.
Two practices operationalize this. Score experiments on information per week rather than on wins: a tight null that closes a question is worth more than a wide win that reopens three. And make the retro question “what does this rule out?” rather than “what went wrong?” — the first produces a durable record, the second produces defensiveness and no record at all.
Ten questions, in deposition order — general foundation first, then the specific attacks, so that a readout which fails early does not waste the room’s time on interpretation.
These are leading questions for a friendly witness. You are not trying to destroy the analyst — you want them to walk in already knowing the ten and having answered them in the document. The goal of a cross-examination checklist that everyone has read is that it never has to be conducted.
This is the complete text of the course. With JavaScript enabled, this same page runs the interactive edition — a self-diagnostic that reorders the syllabus around your gaps, knowledge checks, applied worksheets with model answers, and a spaced-repetition review queue — with progress saved locally in your browser.