Why is the prior-year comparison missing for the first year of each series, which is expected, but also for some interior years where prior data exists? These gaps suggest either a data integrity issue or an intentional exclusion. Should we investigate whether there's a filter being applied that we don't expect?
I suspect the 2016 cohort is being affected by survivorship bias. Should we note that limitation?
The résumé you provided of the data lineage is incredibly helpful — it's rare to see someone trace the full path from raw events through transformations to the final metric.
The outlier in Q2 skews the average so much that the median might be more meaningful. Should we show both? See /u/plg_matteo_nair/p/plot-0001.