# NYC chronic-absenteeism deep dive — plan

*kx-jin, 2026-06-09. Status: Phase 0 complete, Phase 1 starting. Companion doc: [07_absenteeism_metric_alignment.md](07_absenteeism_metric_alignment.md).*

## Context

We met with the Chalkbeat NYC team and aligned on going deep on **one outcome measure first — chronic absenteeism — to validate the overall approach**. Their questions: What are the NYC trends overall? How much school-level variation is there? And which schools are outliers relative to socioeconomically similar peers — in both directions (high-SES schools with high absenteeism, low-SES schools with low absenteeism)?

This doc is the roadmap so you know what's coming. Work iterates locally and lands here as small, single-purpose PRs (same pattern as the `verify/` series). Each phase has a human checkpoint before the next.

## Design decisions — defaults we're proceeding with (flag if you disagree)

1. **Per-grade representation: per-grade metric keys** (`chronic_absenteeism_rate_grade_pk/k/1…12`), following the existing `ela_grade3_proficiency` convention — no change to `school_year_metrics` or its unique constraint. **Scope: all grade rows, all subgroups.** The source file holds 961,914 grade-level cells vs the 171,844 all-grades cells we ingest today. Suppression supports this split: grade-level **All-Students** cells are only **3.9% suppressed** (so per-school grade trajectories and pseudo-cohort tracking — grade g in year y vs grade g+1 in year y+1 — are well-populated), while grade×subgroup cells run 23–50% suppressed (Ethnicity worst); we keep the loader convention of storing suppressed cells as flagged nulls so coverage is documented. Main cost is ~1M new rows and a longer percentile batch job, not code.
2. **Second publisher (NYSED): one canonical `chronic_absenteeism_rate` key with per-universe source routing** — DOE InfoHub for districts 1–32+75, NYSED for charters — mirroring the Philly SDP/FRPI pattern (with year semantics pinned explicitly, avoiding the `graduation_rate_4yr` year-indexing mismatch). Provenance stays per-row via `source_load_id`; the DOE-vs-NYSED comparison on district schools lives in `verify/` as a reconciliation component, not in the serving DB. **Explicit revisit condition:** if that reconciliation shows large or systematic district-school discrepancies, we revisit routing — plausibly promoting NYSED to primary for *all* schools for cross-universe consistency (at the cost of a one-year lag). The NYSED loader will be written so re-routing is a re-run, not a rewrite. This is the one decision that's mildly costly to reverse after ingestion, so it's the most important one to object to *now*.
3. **Verification page: in-app, public, under the methodology surface** (e.g. `/methodology/verify`). v1 renders from committed `verify/` artifacts (`docs/qa_reports/*.json`) + the schools directory — no new tables, no write path. Per-school card: served value → exact source file/sheet/cell reference → Snapshot deep link (`tools.nycenet.edu/snapshot/{year}/{dbn}/{EMS|HS}`) with expected-difference labeling → NYSED report-card link. I'll sketch the page before building it.

## Phases and the PRs to expect

**Phase 0 — Metric alignment (done).** Which metric, which sources, do they agree. Output: the metric-alignment memo (doc 07). Headline = `chronic_absenteeism_rate` (the NYSED definition already ingested); companion = average daily attendance; plus per-grade rates. Trend window 2021-22→2024-25 with 2018-19 baseline; 2019-20/2020-21 excluded as COVID-regime years (matches `reference-scaffold/config/comparability.yaml`).

**Phase 1 — Data foundation.**
- ✅ **PR #9** (open): attendance loader subgroup fix (+ request to re-run with refreshed school directory).
- **PR: ADA + per-grade ingestion** — ingest `% Attendance` and the PK-12 grade rows from the already-registered attendance file, per decision 1 above (per-grade keys, all subgroups).
- **PR: NYSED report-card ingest** — promoted from "validation check" to first-class source, because Chalkbeat wants **all NYC public schools**: the DOE file covers districts 1–32+75 only, while NYSED publishes school-level chronic absenteeism for charters (~281 schools, >15% of students). Scouting confirmed coverage through **2024-25** (SRC2025 released 2026-04-30), so the all-schools universe holds for every analysis year; 2019-20 was never published (and is excluded anyway), and **D79 does not appear in NYSED data** (stays excluded, documented). Concrete path: SRC2019/2022/2024/2025 zips (Access-only, ~1.4 GB; `mdbtools` to export), tables `ACC EM/HS Chronic Absenteeism` keyed on 12-digit BEDS `ENTITY_CD`; BEDS↔DBN is **deterministic for district/D75 schools** (borough-county + CSD + `0001` + school number) and needs a crosswalk only for charters (NYC OD `wg9x-4ke6` has both codes; current LCGMS export for post-2020 charters; `dbn-beds-crosswalk.ts` is currently a stub). Also: charter demographics from NYSED (the DOE Demographic Snapshot is district-only), and quantify the **Success Academy attendance non-reporting** (`DATA_REP_FLAG=N` since 2023-24). Representation per decision 2 above; the district-school overlap between the two publishers is itself the independent (iii-b)-style check NYC has been missing, and it doubles as the gate on the routing decision.

**Phase 2 — Base-data validation + human verification.**
- **PR: verify/ extensions** — components for the new metrics, subgroup coverage beyond `ALL`, the NYSED cross-check, and resolving two open questions: the Snapshot spot-check tail (suspected accountability-cohort denominator) and the 2024+ Snapshot's renamed `val_chronic_absent_*` varnames.
- **PR: in-app verification page** — per decision 3 above. Goal: any Chalkbeat reporter can verify any school in two clicks. Sketch comes before code.

**Phase 3 — Derived metrics (method bake-off).** Four candidate "vs-expectation" measures, compared on list-overlap, year-over-year stability, and SES-correlation diagnostics; then we settle on 1–2:
1. the existing K-NN peer percentile;
2. NYC-official peer averages harvested from the Snapshot API (`cavg_student_chronic_absent`, 2024+ `pavg_chronic_absent_*`);
3. WLS residual-z — lifting `docs/design/reference-scaffold/pipeline/analysis/` per [03_integration_plan.md](03_integration_plan.md) C1 (additive columns, flip-diff vs the decile method before any default flips);
4. a lagged-dependent-variable variant (school-level analogue of Michigan PEER's attendance value-added; their headline caution transfers: raw attendance correlates −0.56 with school disadvantage but value-added correlates **+0.42** — raw rankings flatter advantaged schools).

Plus: per-school trends within the post-COVID regime, and an enrollment-composition decomposition (how much of a school's change is predicted by its demographic change vs residual).

**Phase 4 — Derived-metric validation.** Golden hand-recomputed cases, peer-card face-validity review, sensitivity panel (K, covariates, windows), and a stability gate (outlier labels must persist ≥2 years). Recipe pages in the existing format.

**Phase 5 — Analysis with Chalkbeat.** Descriptives → variance/prediction analysis → named outlier lists both directions, iterated with the Chalkbeat team. Story/recipe surfaces land here as PRs if useful.

## What we'd ask of you

1. Review cadence on the PRs above (all small except the NYSED ingest).
2. After merging ingestion PRs: `prisma db push` + loader re-runs (called out in each PR body), since we deliberately work against a read-only role.
3. A look at the three design decisions at the top — we're proceeding on those defaults, so objections are most useful early (especially decision 2, the publisher routing).
