The four choices every cohort hides
| Choice | Options | How it moves the number |
|---|---|---|
| Anchor | Registration, first deposit, first bet | Registration cohorts include everyone who never funded; FTD cohorts look far healthier on the same players |
| Window | Player-relative (day 0–30 from each anchor) vs calendar (activity in a month) | Calendar windows mix cohort ages; player windows are comparable but mature later |
| Filters | Test, bot, fraud, self-excluded accounts in or out | Funnel and retention rates shift whenever the exclusion list changes — silently, unless the filter set is versioned |
| Censoring | Immature players included, excluded, or projected | Including them understates recent cohorts; the newest campaign always "underperforms" |
None of these choices is wrong. What is wrong is leaving them implicit — because then every team makes them independently, and the monthly meeting spends its hour discovering the differences instead of reading the business.
A worked illustration (synthetic numbers)
Take an illustrative cohort of 1,000 registrations, of whom 400 deposit. Day-30 "retention" can honestly be reported as 12% (active players over all registrations), 30% (active over depositors), or 45% (active over depositors, whale-trimmed medians, mature subset only) — from the same underlying behaviour. All three are defensible; none is comparable to the others. The fix is not choosing the "right" one; it is naming the one you use and never switching mid-chart.
Reading a cohort table without fooling yourself
- Read down the diagonal, not across the bottom row. The newest row is always the most censored; its tail is missing data, not decay.
- Watch the denominators between rows. A retention "improvement" that coincides with an acquisition-mix change is usually the mix, not the product. Split by source before celebrating — media-buy and organic cohorts behave structurally differently.
- Trim or median the value metrics. Averages over stake-distributed populations answer to a handful of accounts; decide the trimming rule once and keep it.
- Anchor product claims to matched cohorts. "Players who used X retain better" is self-selection until the comparison group is matched or randomised — the recurring theme of every measurement page in this Academy.
- Respect the second-deposit checkpoint. FTD-to-second-deposit matures fast and predicts the tail early; it is the honest early read while LTV horizons are still censored.
Where cohorts live in the stack
Cohort analysis is a consumer of everything upstream: the event schema supplies the anchors and both clocks, the warehouse supplies player-grain facts and versioned filters, and the dictionary pins the four choices so that "day-30 retention" means one thing in every deck. When a cohort number surprises you, debug in that order — schema, warehouse, definition — before believing the behaviour changed.
The reporting contract
Any cohort figure that leaves the analytics team should carry its four choices inline: anchor, window, filters, censoring treatment — one line of fine print that costs nothing and ends most disputes before they start. A number without that line is an invitation for someone to reproduce it differently and call yours wrong.
Continue reading: Dashboards by role — where these reads get consumed daily. Cohort retention — the glossary definition this page operationalises.