Perception gapsFull briefing
How much data does NYC suppress, and what does it hide?
The longer form of the briefing for this story: the original thesis, the supporting findings the data team flagged, and the reporting directions a journalist could follow. For the short version and the data analysis, see the main story page.
Thesis
Subgroup data with small N is suppressed for student privacy. Across our metrics, ~7-10% of subgroup cells are suppressed. The suppression is non-random — it falls on the smallest, most-distinctive subgroups. What can we infer about the SCHOOLS where the most data is hidden?
Supporting findings
- Across school_year_metrics, % rows with suppressed=true: chronic absenteeism 7%, test scores 22% (test results have more subgroup cells).
- Suppression hits SWD, ELL, race subgroup cells hardest at schools where those subgroups are small.
- Schools with the most suppressed cells are typically the most demographically-homogeneous (since the small subgroups have small N).
- Information loss: when we can't see a subgroup, we can't see how the school is serving them.
- Identify the schools where most subgroup cells are suppressed — they're the schools where 'the average looks fine' but individual subgroup outcomes are hidden.
- Is the suppression bar (n<5? n<10?) appropriate? Federal guidance varies.
- Compare with NYS's reporting: do they suppress at the same threshold?
- What can be done? Multi-year rollups can reduce suppression (combine 3 years of small cohorts).
- Aggregated borough/district-level reporting can fill in gaps.
- Frame: 'Privacy is good. But privacy without context becomes invisibility for the kids who most need to be seen.'
Reporting directions
- Compute suppression rate per school per subgroup.
- Identify the schools where 50%+ of subgroup cells are suppressed.
- Compare with NYS, US Ed Dept guidelines on small-N reporting.
- Profile the kids hidden by suppression — interview an ELL coordinator at a school with 4 ELL kids and a suppressed cell.