← Methodology

Data quality

Every number on this site comes from a public file published by NYC DOE, NYC Open Data, or NYSED. This section documents how we check that what we serve matches what was published — and what we found when the publishers disagree with each other.

14 of 14

validation reports exact

0 value mismatches across 868,533 served cells re-read from the source files (base case ii).

22

findings on the ledger

8 verified · 5 fixed · 3 open · 6 in progress.

Metrics

Chronic absenteeism: methodology and verification

The definition we use, the three datasets that carry it, key data limitations, and every check we ran. Includes the cross-publisher verification: 10 of 12NYSED-vs-DOE panels reconcile school-by-school at the median; the two that don't (2024-25 grades 1–8 and 2022-23 grades 9–12) — plus a COVID-era high-school tail that passing medians hide — are documented with what is and isn't known about why.

  • ela all proficiencybase case: 6,565 cells, 0 mismatches — PASS · snapshot context: 73% within ±1.5ppevidence · full writeup pending
  • ela grade3 proficiencybase case: 4,646 cells, 0 mismatches — PASSevidence · full writeup pending
  • ela grade7 proficiencybase case: 2,757 cells, 0 mismatches — PASSevidence · full writeup pending
  • graduation rate 4yrbase case: 3,736 cells, 0 mismatches — PASSevidence · full writeup pending
  • math all proficiencybase case: 6,449 cells, 0 mismatches — PASS · snapshot context: 62% within ±1.5ppevidence · full writeup pending
  • math grade3 proficiencybase case: 4,643 cells, 0 mismatches — PASSevidence · full writeup pending
  • math grade7 proficiencybase case: 2,748 cells, 0 mismatches — PASSevidence · full writeup pending
  • survey student 2021-22base case: 73,241 cells, 0 mismatches — PASSevidence · full writeup pending
  • survey student 2022-23base case: 81,849 cells, 0 mismatches — PASSevidence · full writeup pending
  • survey student 2023-24base case: 76,561 cells, 0 mismatches — PASSevidence · full writeup pending
  • survey teacher 2021-22base case: 227,253 cells, 0 mismatches — PASSevidence · full writeup pending
  • survey teacher 2022-23base case: 185,084 cells, 0 mismatches — PASSevidence · full writeup pending
  • survey teacher 2023-24base case: 182,366 cells, 0 mismatches — PASSevidence · full writeup pending

Tools & evidence

How the checks work

  • (i) Completeness: are years or schools missing? Every gap between source and served is counted and explained (or flagged open).
  • (ii) Correctness — base case: does the database match an independent re-read of the exact ingested file? The binding gate.
  • (iii) Correctness — spot check: does it match the public per-school report (same publisher)? Context, not a gate. (iii-b): a genuinely different publisher.
  • (iv) Correctness — computed values: are derived metrics computed correctly from validated inputs? Not done yet for NYC.

The base case (ii) is binding; same-publisher spot checks (iii) are context; when two independent publishers disagree (iii-b), neither is "corrected" toward the other — the disagreement is documented and affected values carry source attribution. Full methodology: verify/METHODOLOGY.md.