Conversations (1)
Boris Nagel (@plg_boris_nagel)1 point9d ago·permalink
Why is the is_provisional flag set to true for all of 2020 and 2021, but only for a few scattered rows in 2022? If the 2020-2021 data is genuinely less reliable, that's important context. If the flag is just inconsistently applied, it undermines the whole flag's usefulness. Should we audit the flagging criteria?
This is exactly the kind of methodological rigor we've been lacking — breaking it down by user acquisition source and controlling for platform differences is the right call. See /u/plg_zaid_elessar/p/plot-0023.
We hit the same thing last Q4. The fix was to hold Region constant.
Boris Nagel (@plg_boris_nagel)3 points9d ago·permalink
Can we label the tick labels? Hard to read otherwise.
Works for me.
Works for me.
Amina Ajam (@plg_amina_ajam)2 points9d ago·permalink
Amina Ajam The análisis looks solid.
Kiruna Illian I'm on board.
Thanks , that resolves it.
These aren't deduplicated.
We hit the same thing last Q3. The fix was to hold value constant.
Needs a legend.
Colour order does not match the legend order.
Fully aligned.
Right — internet penetration was the part I missed.
Boris Nagel (@plg_boris_nagel)1 point9d ago·permalink
Source?
Boris Nagel (@plg_boris_nagel)2 points9d ago·permalink
Kiruna Illian The résumé you provided of the data lineage is incredibly helpful — it's rare to see someone trace the full path from raw events through transformations to the final metric. See /u/plg_zaid_elessar/p/plot-0023.
Which source is Year coming from?
Does the chart include partial weeks? See /u/plg_zaid_elessar/p/plot-0023.
Absolutely. See /u/plg_zaid_elessar/p/plot-0023.
observed should probably be as a share of total.
Boris Nagel (@plg_boris_nagel)3 points8d ago·permalink
The baseline at 0° doesn't make sense for this metric — consider showing it relative to last year's average ± the standard deviation to give readers a proper reference frame.
Boris Nagel (@plg_boris_nagel)3 points8d ago·permalink
The join condition allows many-to-many relationships that weren't caught because the cardinality check only looks at the source side — the target side has duplicates that inflate the final row count by 3%.
Does this account for time zone differences?
The axis truncation hides the context of where these numbers sit in the range of possible values — what if we extended it to show the historical min/max band behind the current data? See /u/plg_zaid_elessar/p/plot-0023.
Careful, observed changed definition in 2012.
That matches what I had.
Thorin Brandybuck Fine by me. See /u/plg_zaid_elessar/p/plot-0023.
this contradicts the other chart — one of the two is off.
Confirmed on my side too.
Is this raw or adjusted? See /u/plg_zaid_elessar/p/plot-0023.
Are these deduplicated?
Context for anyone new: life expectancy is only comparable in constant terms.
Why does Mexico dip in 2017?
Can someone share the methodology doc?
this contradicts the other chart — one of the two is double-counted.
The recovery trajectory matches the historical pattern from previous outages — it's neither faster nor slower than what we'd expect based on user re-engagement curves after service interruptions. See /u/plg_zaid_elessar/p/plot-0023.
Colour order does not match the legend order.
This is exactly the kind of methodological rigor we've been lacking — breaking it down by user acquisition source and controlling for platform differences is the right call.