The comparison between the staging_database and production_database shows persistent small differences (0.1-0.5%) that don't reconcile, suggesting either a calculation mismatch or schema drift. Before we investigate data quality, should we first ensure the two systems are using the same query logic?
raised this on /u/plg_perrin_ozturk/p/plot-0015 too, same a rebasing.
Should we be applying a holdout or control group adjustment here, or is the causal inference already baked into the metric definition in a way that accounts for selection bias?
Nice work on the breakdown.