How much data does NYC suppress, and what does it hide?
A reproducibility log: the data this analysis touched, the queries it ran, the outside sources it leaned on, and the steps a third party would follow to re-run it.
Summary
Counted the share of school_year_metrics rows where suppressed=true, broken out by subgroup. Highlighted the impact on subgroup-gap stories.
Data sources
- Database tableschool_year_metrics
Long-format per-school per-year metric facts (school_dbn, year, metric_key, subgroup, value, suppressed). Loaded from DOE/NYSED public files via scripts/loaders/*.
Steps
SELECT subgroup, COUNT(*), COUNT(*) FILTER (WHERE suppressed=true) FROM school_year_metrics GROUP BY subgroup; compute suppression rates.
Caveats
Suppression rules vary by year and by metric; the share is approximate. Suppression is genuinely meant to protect small-n student privacy, not to obscure outcomes.
Reproduce
Clone the repo, set DATABASE_URL to a Postgres with the project schema loaded, run `npx tsx scripts/loaders/<source>.ts` for any not-yet-loaded data, then issue the queries in the Steps section. The story page also lists the exact `metric_key`/`subgroup`/`year` filters used. Searchable by the answer's headline number — every figure is recomputable from the queries shown.
The recipe lives at data/stories/recipes.ts in the repo. Corrections welcome.