The four choices every cohort hides

ChoiceOptionsHow it moves the number
AnchorRegistration, first deposit, first betRegistration cohorts include everyone who never funded; FTD cohorts look far healthier on the same players
WindowPlayer-relative (day 0–30 from each anchor) vs calendar (activity in a month)Calendar windows mix cohort ages; player windows are comparable but mature later
FiltersTest, bot, fraud, self-excluded accounts in or outFunnel and retention rates shift whenever the exclusion list changes — silently, unless the filter set is versioned
CensoringImmature players included, excluded, or projectedIncluding them understates recent cohorts; the newest campaign always "underperforms"

None of these choices is wrong. What is wrong is leaving them implicit — because then every team makes them independently, and the monthly meeting spends its hour discovering the differences instead of reading the business.

A worked illustration (synthetic numbers)

Take an illustrative cohort of 1,000 registrations, of whom 400 deposit. Day-30 "retention" can honestly be reported as 12% (active players over all registrations), 30% (active over depositors), or 45% (active over depositors, whale-trimmed medians, mature subset only) — from the same underlying behaviour. All three are defensible; none is comparable to the others. The fix is not choosing the "right" one; it is naming the one you use and never switching mid-chart.

Reading a cohort table without fooling yourself

  • Read down the diagonal, not across the bottom row. The newest row is always the most censored; its tail is missing data, not decay.
  • Watch the denominators between rows. A retention "improvement" that coincides with an acquisition-mix change is usually the mix, not the product. Split by source before celebrating — media-buy and organic cohorts behave structurally differently.
  • Trim or median the value metrics. Averages over stake-distributed populations answer to a handful of accounts; decide the trimming rule once and keep it.
  • Anchor product claims to matched cohorts. "Players who used X retain better" is self-selection until the comparison group is matched or randomised — the recurring theme of every measurement page in this Academy.
  • Respect the second-deposit checkpoint. FTD-to-second-deposit matures fast and predicts the tail early; it is the honest early read while LTV horizons are still censored.

Where cohorts live in the stack

Cohort analysis is a consumer of everything upstream: the event schema supplies the anchors and both clocks, the warehouse supplies player-grain facts and versioned filters, and the dictionary pins the four choices so that "day-30 retention" means one thing in every deck. When a cohort number surprises you, debug in that order — schema, warehouse, definition — before believing the behaviour changed.

The reporting contract

Any cohort figure that leaves the analytics team should carry its four choices inline: anchor, window, filters, censoring treatment — one line of fine print that costs nothing and ends most disputes before they start. A number without that line is an invitation for someone to reproduce it differently and call yours wrong.

Continue reading: Dashboards by role — where these reads get consumed daily. Cohort retention — the glossary definition this page operationalises.