Data quality
Every number on this site comes from a public file published by NYC DOE, NYC Open Data, or NYSED. This section documents how we check that what we serve matches what was published — and what we found when the publishers disagree with each other.
14 of 14
validation reports exact
0 value mismatches across 868,533 served cells re-read from the source files (base case ii).
22
findings on the ledger
8 verified · 5 fixed · 3 open · 6 in progress.
Metrics
The definition we use, the three datasets that carry it, key data limitations, and every check we ran. Includes the cross-publisher verification: 10 of 12NYSED-vs-DOE panels reconcile school-by-school at the median; the two that don't (2024-25 grades 1–8 and 2022-23 grades 9–12) — plus a COVID-era high-school tail that passing medians hide — are documented with what is and isn't known about why.
- ela all proficiencybase case: 6,565 cells, 0 mismatches — PASS · snapshot context: 73% within ±1.5ppevidence · full writeup pending
- ela grade3 proficiencybase case: 4,646 cells, 0 mismatches — PASSevidence · full writeup pending
- ela grade7 proficiencybase case: 2,757 cells, 0 mismatches — PASSevidence · full writeup pending
- graduation rate 4yrbase case: 3,736 cells, 0 mismatches — PASSevidence · full writeup pending
- math all proficiencybase case: 6,449 cells, 0 mismatches — PASS · snapshot context: 62% within ±1.5ppevidence · full writeup pending
- math grade3 proficiencybase case: 4,643 cells, 0 mismatches — PASSevidence · full writeup pending
- math grade7 proficiencybase case: 2,748 cells, 0 mismatches — PASSevidence · full writeup pending
- survey student 2021-22base case: 73,241 cells, 0 mismatches — PASSevidence · full writeup pending
- survey student 2022-23base case: 81,849 cells, 0 mismatches — PASSevidence · full writeup pending
- survey student 2023-24base case: 76,561 cells, 0 mismatches — PASSevidence · full writeup pending
- survey teacher 2021-22base case: 227,253 cells, 0 mismatches — PASSevidence · full writeup pending
- survey teacher 2022-23base case: 185,084 cells, 0 mismatches — PASSevidence · full writeup pending
- survey teacher 2023-24base case: 182,366 cells, 0 mismatches — PASSevidence · full writeup pending
Tools & evidence
Verify a school's number
Any school, year, and metric: our value, the exact source row to check, and honestly-labeled cross-references. Covers the attendance metrics today; expands metric by metric.
Findings ledger
Everything we have established, one entry per finding — verified, fixed, open, or in progress. Never deleted, only updated.
Working analysis (not verified)
The descriptive analyses in progress — trends, grade patterns, distributions — clearly banner-marked pending editorial review.
All validation reports
The per-metric machine-generated reports: cells checked, mismatches, agreement rates, links to full write-ups.
How the checks work
- (i) Completeness: are years or schools missing? Every gap between source and served is counted and explained (or flagged open).
- (ii) Correctness — base case: does the database match an independent re-read of the exact ingested file? The binding gate.
- (iii) Correctness — spot check: does it match the public per-school report (same publisher)? Context, not a gate. (iii-b): a genuinely different publisher.
- (iv) Correctness — computed values: are derived metrics computed correctly from validated inputs? Not done yet for NYC.
The base case (ii) is binding; same-publisher spot checks (iii) are context; when two independent publishers disagree (iii-b), neither is "corrected" toward the other — the disagreement is documented and affected values carry source attribution. Full methodology: verify/METHODOLOGY.md.