How to check the growth narrative for confirmation bias
Confirmation bias is the most persistent distortion in growth reporting. It shows up when an analyst selects metrics that support the interpretation they already prefer, or when a flat primary metric gets ignored because a secondary metric nudged up. The reviewer's job is to compare what the analyst concluded against the full metric set, not just the metrics the analyst chose to highlight.
Check whether the primary metric was declared before the analysis window opened. If the analyst switched from revenue to click-through rate after revenue came in flat, that is a retrofitted narrative. Also check whether the analyst is only reporting segments that won while ignoring the segments that lost. GrowthLayer's audit of 57 claimed wins found only 17 survived rigorous standards. That 70% gap is confirmation bias compounded over time.
- Compare conclusion to the full metric set. Not just metrics the analyst chose to feature.
- Verify primary metric was declared before the window opened. Swapping after is retrofitted narrative.
- Flag any report where a flat primary gets silently replaced by a secondary that moved.
- Check every segment, not just the winning one. A +14% mobile win hidden inside a flat average is a finding.
How to detect survivorship bias in the evidence set
Survivorship bias occurs when the analyst builds conclusions only from tests that won or channels that worked, ignoring the tests that lost and the channels that failed. A program that only publishes winners is signalling it is hiding data. Real programs publish 60-70% non-winners because most experiments are flat or negative. A growth narrative sourced entirely from winning tests is sampling from the top of an incomplete distribution.
The reviewer should demand the full test log: all experiments run in the period, not just shipped winners. Check whether flat or losing tests produced insights that should constrain the recommendation. If an interpretation cites a +15% test win without acknowledging that three previous tests on the same surface lost or were flat, the confidence in that interpretation should be downgraded.
- Demand the full test log. Not just shipped winners. Real programs publish 60-70% non-winners.
- Flag interpretation built only from winning tests while ignoring lost and flat experiments.
- Ask what flat and losing tests in the same period tell you about the channel or surface.
- Downgrade confidence when a +15% win ignores three prior flat tests on the same surface.
How to spot segmentation bias hiding in overall averages
Segmentation bias is the most underrated distortion in growth interpretation. An overall result that reads as neutral or flat often hides a large segment-level effect. A test can be up 14% on mobile and down 4% on desktop and still show a flat aggregate. The business decision changes completely if you see what the average hid. Pre-declare segments before the analysis, not after.
Running ten segment splits on a flat result and reporting only the one that won is p-hacking. The false-positive rate on post-hoc segment analysis runs 30-50% depending on how many cuts you make. Require the analyst to declare analysis segments before the data is reviewed. Any segment finding discovered after unplanned exploration must be treated as a hypothesis for the next test, not as a conclusion for this one.
- Split every overall result by device, source, and new vs returning before accepting it.
- Pre-declare analysis segments before data is reviewed. Post-hoc splits are hypotheses, not findings.
- Treat a +14% mobile signal inside a flat aggregate as a separate decision branch.
- Flag any report that ran ten segment cuts and published only the one that looked significant.
How to test whether the interpretation survives pre-registration
Pre-registration is the single highest-leverage anti-bias intervention in growth analysis. Before any test result reaches review, the analyst must have written down the primary metric, the minimum sample size, the planned duration, and the stopping rule. No result can be called a win if it does not meet these pre-registered criteria regardless of how the numbers look. This discipline cuts the false-positive rate by three to four times in the experimentation literature.
The reviewer should also apply the Winner's Curse discount. A +15% test win is realistically +9-10% in production because the reported lift is selection-biased upward. The industry shrinkage factor is 20-50% for individual wins and 30-50% for portfolio aggregates. Growth recommendations that report raw test lifts without a shrinkage adjustment are systematically overstating impact. Require both numbers: raw lift and shrinkage-adjusted estimate.
- Require pre-registered primary metric, sample size, duration, and stopping rule before every test.
- Reject any result called a win that doesn't meet pre-registered criteria regardless of how the numbers look.
- Apply a 30-50% shrinkage discount to headline test lifts before briefing stakeholders.
- Report both raw observed lift and shrinkage-adjusted estimate. Raw is ceiling. Adjusted is forecast.
When to approve a growth interpretation versus challenge it
Approve when the interpretation passes four gates: primary metric was pre-declared and held, full test log was examined not just winners, segments were pre-declared and all were reported, and shrinkage adjustment was applied to headline lifts. The interpretation should name what the program learned from flat and negative results alongside what it learned from wins. An interpretation that explains both what worked and what failed is more credible than one that only explains success.
Hold when the interpretation cherry-picks metrics, segments, or tests. A report that switches primary metrics mid-read, runs undeclared segment splits, or cites only winning experiments without acknowledging the full log should be sent back. Also hold if the headline lift is reported without shrinkage and the recommendation implies the raw number will replicate in production. Tell the analyst exactly which gate failed and what a passing interpretation looks like. The reviewer is not saying the work is wrong. The reviewer is saying the work is not yet sufficiently defended against bias.
- Approve when pre-declared metric held, full log examined, segments reported, shrinkage applied.
- Require interpretation to explain what flat and negative results taught alongside what wins taught.
- Hold for metric-switching, undeclared segment fishing, or winner-only test logs regardless of headline lift.
- Send back any raw lift reported as production forecast without 30-50% shrinkage adjustment applied.
Sample Review Note
All five gates checked. The analyst pre-declared revenue per session as the primary metric and held it through the full read. The test log includes three flat experiments and one loss alongside the two wins cited in the recommendation, and the analyst explains how the flat tests constrained the scope of the recommendation. Segments were pre-declared as mobile, desktop, new, and returning, and all four segments were reported regardless of result. The headline +12% lift is reported alongside a shrinkage-adjusted estimate of +8%, derived from the program's historical 33% shrinkage factor. Interpretation approved with quarterly re-audit to verify replication.
Recheck triggers: if post-rollout revenue per session is below the shrinkage-adjusted estimate in the first full quarter, or if any guardrail metric degrades beyond its defined threshold. The next quarter's bias audit should re-verify that the full test log, not just winners, continues to be published. Decision stays approval-gated until reviewer accepts the interpretation as defended against the five bias gates.