Method bake-off — results memo (DRAFT)
Executive summary
The question
We are interested in understanding which NYC schools stand out on chronic absenteeism — in both directions: schools with much higher absenteeism than you'd expect given their students, and the encouraging cases, with much lower absenteeism than their high-need student bodies would predict — plus which schools are improving or backsliding over time.
The raw chronic-absenteeism rate can't answer this. At the school level it largely tracks how economically disadvantaged the students are (correlation ≈ 0.57): rank schools by the raw rate and you mostly re-rank them by poverty. Independent Michigan evidence shows the same reversal: a Detroit PEER study (Singer, Lenhoff & Lyle, 2026, statewide Michigan) found a school's raw attendance rate correlates −0.56 with its students' economic disadvantage, but the school's attendance value-added correlates +0.42 with disadvantage (Figs. 3–4, both p<0.001) — adjustment flips the sign, and raw rates flatter advantaged schools. So we need a "vs-expectation" measure: how a school compares to what you'd expect given its student body — both for where it stands today (status) and whether it's moving (trend) — without simply reproducing the poverty map.
The approaches we compared
Four candidate measures, each a different way to define "expectation":
- (a) Peer-group percentile (the method already live on the site) — compare each school to its ~40 most demographically similar schools and rank within that group.
- (b) NYC's own published benchmark — reuse the comparison-group average the city already publishes in its Snapshot tool.
- (c) Statistical model (regression residual, "residual-z") — predict each school's expected absenteeism from its student body and structure, then read the gap between actual and expected. Included (the factors we definitely adjust for): poverty/economic need, homelessness, English-learner and special-education shares, school size, grade range, admission type, and grade mix. Under consideration: race composition — run both with and without it, since whether to include race in a published "expected absenteeism" model is an editorial, not statistical, decision (left to the checkpoint). Held out: borough — geography isn't a student-body trait, so it stays out of the model and is used as an orthogonality check instead (so the measure can't simply be flagging where a school is).
- (d) Beat-your-own-track-record — the same model plus the school's prior-year absenteeism, asking whether it did better or worse than its own history predicts.
How we judged them — and why each test
We fixed five tests before computing anything, so the results couldn't quietly reshape the criteria. In priority order:
- Stays stable year to year. An outlier label should reflect a real, persistent feature of a school, not one year's noise — otherwise a single-year flag can't be trusted. We benchmark this against how stable the raw rate itself is.
- Actually removes student-body factors — the whole point. The decisive version of this test: does the measure neutralize factors it was never explicitly given (student homelessness, school size, borough)? A measure can look poverty-neutral while quietly smuggling demographics back in; this is what catches that.
- Do the methods agree on who's flagged? If the choice of method swings roughly half of who gets named, that's a finding in itself — and something to disclose, not hide.
- Matches NYC's official benchmark. A convergence check: our expectation should broadly track the city's own published comparison numbers; a large divergence is a flag to investigate.
- Explainable in one sentence. Whatever we publish has to be statable to a reader without a footnote. This can break ties — but only ties.
A separate gate governs any named school: it must be an outlier in at least two of the three most recent years, with extra caution for small schools, where a handful of students can swing the rate.
What we recommend
- Status (where a school stands): the statistical model, candidate (c). It is the only approach that genuinely strips out student-body factors — including the ones it was never handed — while costing only a little year-to-year stability versus the peer method (≈0.83 vs ≈0.84, against a raw-rate baseline of ≈0.93). NYC's own published benchmark independently agrees most closely with (c) in the cleanest year.
- Show it to readers through the peer-group percentile (a). It is the one you can say in a sentence — "among the 40 most similar schools, this one had the Nth-highest absenteeism" — so it stays as the reader-facing display, but not as the standalone analytical measure: it still carries traces (correlations of 0.2–0.3) of factors it never compared on.
- Trend (is it changing): the change in the model's gap across the last three years, plus recovery versus pre-pandemic (2018-19).
- Dropped: the beat-your-own-track-record measure (too volatile to anchor a "where it stands" claim), and NYC's published benchmark as our source of truth — which, given its real appeal, is worth spelling out below.
Why not just reuse NYC's benchmark?
Leaning on the city's own comparison number is genuinely appealing: it would be more explainable to readers, and it would spare us from having to defend a method of our own. We looked hard at it — and it can't carry the analysis.
First, what NYC actually publishes for chronic absenteeism on the school Snapshot:
| School year | Chronic absenteeism reported? | School-specific comparison group? |
|---|---|---|
| 2022-23 and earlier | Yes — directly ("Students chronically absent: X%") | No |
| 2023-24 onward | Yes — as the complement ("Students with ≥90% attendance"; chronic = 100 − X) | Yes ("Comp group*") |
So the chronic-absenteeism rate itself is available every standard year — NYC only flips how it's framed in 2023-24 (from the rate to its attendance complement) — but a per-school comparison benchmark exists only from 2023-24 on. (COVID years 2019-20/2020-21 are sparse or non-reported.)
That second column is the crux — and it's why reuse can't work, for reasons worth stating plainly:
- There is no reusable peer group to inherit. NYC has never published a list of comparison schools you could show a reader. Through 2022 its "Comparison Group" was a student-level matching construct — ~50 statistically matched students per student, pooled to the school level — and from report year 2023 on NYC switched to an MIT Blueprint Labs regression, a model-based counterfactual estimating "how this school's students would have performed had they instead attended a random NYC school" — with no group of schools at all. (NYC's dashboard note adds that "for years before 2023, the comparisons are based on a student-matching method.") So "reuse what they have" isn't actually available: there is no transparent grouping to adopt, and NYC's current benchmark is itself a regression model — no more transparent, or easier to explain to a reader, than ours.
- Even the genuine school-specific benchmark covers only part of what we need. NYC does publish a school-specific comparison value for chronic absence (the Snapshot
pavgfield), and we use it — but only for 2023-24 and 2024-25, not the third scored year (2022-23) or the 2018-19 pre-pandemic baseline the trend rests on; it comes on a single grade-band, integer-rounded basis (cleanly usable only for single-band district schools); and the methodology changed mid-window. Because it is a black-box model we cannot recompute, we also cannot extend it to the years or schools it omits. So it can serve as a check on part of the analysis, but not as the consistent, all-years, all-schools spine the question needs. - And where NYC's school-specific benchmark does exist, our model already agrees with it — most closely of all four candidates (correlation ≈ 0.84 in the cleaner year). So we lose nothing by using our own: we reproduce NYC's logic with a transparent, every-year, all-schools measure we control and can document, and we keep NYC's
pavgin the analysis as an external validation check (criterion 4 / §9) — just not as the measure itself.
An independent re-derivation of every selection-driving number (2026-06-17) confirmed this recommendation. The sections below give the full method and every figure.
Candidate schools & decisions for the checkpoint
The candidate schools: named lists are in the outlier candidates view, with each name's robustness to the grade-control choice in the grade-control sensitivity view. These are fine to keep on the working draft site (it's unadvertised/noindex, and the existing public view already surfaces outliers); they're preliminary leads to investigate, not verdicts, and not for an advertised or published story until validated and reviewed (decision 3). Headline sets (candidate c, spec F): 9 "above expectation" / 18 "below expectation" schools double-confirmed by our model and NYC's comparison group in both 2023-24 and 2024-25 (the most defensible), and 20 / 18 flagged on all five candidate measures. Trend candidates — schools moving fastest vs expectation — are in the separate trend candidates view: on the strong tier (|Δz| ≥ 2 across 2022-23 → 2024-25, race-unaware), 17 worsening / 14 improving, of which 16 / 13 are "sustained" (the move registers at both the 2023-24 and 2024-25 endpoints — a dual-endpoint persistence gate, the change-space analog of D10 — screening out one-year blips).
Decisions / discussion for the checkpoint — five open calls:
1. Race in the model? (editorial — D3). The model runs with or without the race-composition covariates. Without race, the measure asks "vs students of similar poverty, language, disability and need"; with race, it also conditions on racial composition. Both pass the orthogonality checks and rank-correlate 0.94–0.97 (≈25 schools/yr/direction swap), so the choice moves the tails, not the core — but it's an editorial, not statistical, call: a public "expected absenteeism" model that bakes in race can read as setting different expectations by race. We report both; the team decides what to publish (e.g., publish race-blind and show race as a separate lens).
2. Composition-stability threshold (D11) — still open. D11 flags schools whose student mix barely shifted over the window, so an absenteeism change is more plausibly the school's own doing rather than a demographic shift. "Barely shifted" can be defined strictly (max drift across all years below threshold → only 3.0% of schools pass — a tight narrative-priority filter) or loosely (year-over-year drift → 27.0% pass). It doesn't change who is an outlier — only which outliers earn the "composition held still" corroboration. Pick one definition, or carry both.
3. Naming policy + validation gate. The candidate schools can stay on the working draft site now (it's unadvertised, and the existing public view already shows outliers). The gate is for an advertised or published story: (a) which set to feature — the most defensible is the robust E∩F / double-confirmed schools (flagged by our model and NYC's comparison group, both years), not the spec-dependent F-only ones; (b) Phase-4 validation first — golden hand-recomputation of sample schools, per-school face-validity, and confirming each is not a data artifact (e.g., the 2024-25 EM publisher split [A.11], the COVID-era HS tail [A.20], suppression); (c) the F-only schools (4 worse / 5 better) reviewed individually, since they hinge on the grade-control choice. Confirm the bar for story-publication.
4. Scope — charters. These lists are NYC district schools only: the peer-adjusted model needs demographics, and the DOE Demographic Snapshot is district-only (charter demographics from NYSED is a named follow-on). Charters are ~15%+ of students, so this is a real coverage gap. Decide whether to ship district-only for a first story and add charters later, or wait for charter coverage. (Charter raw CA exists via NYSED for descriptive context — just not the adjusted measure.)
5. Trend framing. For "is a school improving or sliding," we recommend reporting the direction + magnitude of the school's (c)-z change across the three scored years (2022-23 → 2024-25), plus its recovery vs 2018-19 — not a fitted slope with a p-value, because three scored points cannot support a significance test. To separate genuine multi-year movement from one-year blips, the trend candidates view applies a persistence gate: a school counts as sustained only if its change registers at both the 2023-24 and 2024-25 endpoints (the change-space analog of the locked D10 dual-endpoint rule) — blips stay visible but tagged, and each school's composition drift is shown alongside, since a changed student body is the main innocent explanation for an apparent trend. Confirm this framing (vs wanting a formal trend test the data can't bear).
6. Trend expectation basis — within-year vs pooled (open, low-stakes). The trend z is currently standardized within each year (each year's model is refit), so a school's Δz blends its own movement with the model being re-drawn annually: as the city recovers, the expectation bar both drops (~1.5pp/yr) and compresses (the highest-need schools are predicted ~2.6pp lower in 2024-25 vs ~1pp for the lowest-need), so a school standing still can show a rising z even with a flat raw rate. The alternative is a pooled model with year fixed effects — one set of coefficients across 2022-23 → 2024-25, with year dummies absorbing the citywide level — so a school's Δz reflects only its own movement against a stable expectation structure. A quick diagnostic (race-unaware, pooled_trend_diagnostic.py) shows the two agree almost entirely: Δz rank-correlation 0.97, and of the strong-tier names 13 of 17 "worsening" and 13 of 14 "improving" are common; the pooled model is slightly more conservative on the worsening side (drops 4, all marginal — including the one-year blip 19K502 and a school with ~4× enrollment drift). So this is a robustness footnote, not a story-changer. Recommendation: keep within-year z as primary (it answers "vs this year's peers," the cleaner peer-relative claim) and report pooled+year-FE as a concurring robustness check; switch only if an editorial preference for a fixed, easier-to-communicate expectation outweighs the peer-relative framing.
Additional details
Status: DRAFT, gated preliminary (D9). Results of the bake-off specified in
docs/analysis/absenteeism/03_bakeoff_design.md(design lock, D1–D12). Every number below is traceable to a named file in this workspace (absenteeism-deep-dive/bakeoff/):evaluation.json,candidate_c_stats.json,candidate_d_meta.json,analysis_table_meta.json,peer_groups/peer_groups_meta.json,recovery_meta.json+recovery.csv,qualifiers_report.txt,flipdiffs.csv,gated_shortlists.json,k_sensitivity.json. No named-school lists appear in this memo (D9):gated_shortlists.jsonholds the names; this memo reports counts and characterizations only. No DB access was used to write §1–8 (§9 and the K-sensitivity artifact each used one SELECT-only extract, noted in place). Independent verification (2026-06-17): every selection-driving number was re-derived from raw inputs via code paths independent ofevaluate.py/engine.py(candidate (c) reproduced to machine precision, worst |Δz| = 5×10⁻¹⁴; K=40 peer groups reproduced 1442/1442). The recommendation is confirmed; the corrections that pass surfaced are applied below — full change log inVERIFICATION_REPORT.md.
2026-06-11 · computed outputs dated 2026-06-11 · repo read-only for this draft.
Reading the codes — the "D#" design decisions
The bake-off was pre-registered in the design memo (03_bakeoff_design.md), which locked twelve decisions (D1–D12) before any result existed. The sections below reference them by number; in plain terms:
- D1 — four candidate measures enter; we exit with 1–2 (one for status, one for trend).
- D2 — candidate definitions were frozen up front; any change is logged as a dated amendment, never made silently.
- D3 — the model's covariate set is frozen; whether to include race composition is an open editorial call (we run it both ways).
- D4 — "expectation" is computed per year for status, and as a trend of year-specific residuals for trend.
- D6 — universe & window: NYC district schools, post-COVID; charters excluded for now; scored years 2022-23 → 2024-25.
- D7 — the five evaluation criteria and their priority order were fixed before any scoring.
- D8 — the bar for naming a school: an outlier in ≥2 of the 3 most recent years, with extra caution for small schools.
- D9 — named-school lists stay gated/preliminary until the Chalkbeat checkpoint (this memo reports counts only).
- D10 — recovery metrics measured against the 2018-19 pre-pandemic baseline.
- D11 — composition-stability screen: flag schools whose student mix barely shifted (so a change is more plausibly the school's own doing).
- D12 — cohort-coherence score: do a school's grade cohorts move together (a corroborating signal).
1. Results summary
Amendment A1 — grade-composition control, spec F (2026-06-18, per D2/D3). Candidate (c)'s school-wide residual was found to still carry grade-mix signal (corr +0.11–0.13 with grade-mix-expected CA — ~4× the criterion-2 bar; PK-heavy elementaries and grade-12-heavy high schools tilted "worse"). The covariate set was amended: (i) added
gmx— an indirect-standardized grade-mix-expected-CA index (spec F: Σ citywide grade-rate × school grade-share) — which drives grade-mix leakage to ~0 (corr +0.01/+0.03/+0.03); (ii) droppedboroughfrom the model to a held-out criterion-2 check, matching the frozen §3 covariate set (its earlier inclusion was an unlogged deviation).log_enrollmentis kept (a with/without sensitivity,logenr_sensitivity.json, shows omitting it leaves a −0.06 to −0.08 residual size correlation). E-vs-F decision + full rationale:grade_control_decision.md. All §1–§9 figures below reflect this F re-run (candidate (c)/(d) refit;evaluation.json, gated lists, §9 recomputed;candidate_*_preF.csvretain the prior fit). The headline recommendation is unchanged — (c) for status; the named lists shift (see the grade-control sensitivity view: E∩F = robust, F-only = review).
The pre-committed D7 criteria select candidate (c), the WLS residual-z, as the status measure: it is the only adjusted candidate that passes the SES-orthogonality test against omitted factors (in-model SES/structure factors |r| ≈ 0 and grade mix now controlled (corr ≈ 0 with grade-mix-expected CA, down from 0.11–0.13; A1); the genuine held-out checks pass — race-share corr ≤ 0.13 (without-race) and borough R² ≤ 0.033 — per-factor detail in §2) while costing about the same stability as the K-NN percentile (YoY Pearson 0.826/0.840 vs 0.837/0.858, against a raw-CA baseline of 0.928/0.936 on the same schools; evaluation.json criteria 1–2). Candidate (a), the live K-NN peer percentile, retains material correlation with factors its distance never saw — STH share 0.23–0.32, ENI 0.20–0.23, log enrollment −0.22 — so it fails the criterion-2 test as a standalone status measure, though it wins the criterion-5 one-sentence test and remains the natural display/communication companion. Candidate (d), the lagged-DV residual, fails criterion 1 outright as a status measure (YoY Pearson 0.02 and −0.12) and per the D7 order is not rescued; its D8-persistent counts (47–55) sit at the independence chance baseline (~39); it can only be read as a year-shock/growth measure. Candidate (b) could not adjudicate: the ingested cavg benchmark is degenerate — 1–2 distinct values per report-year × band — so criterion 4 was downgraded to a level-calibration check, which all candidates pass within the benchmark's own basis offset. The with/without-race variants of (c) rank-correlate at 0.94–0.97 and swap roughly 25 schools per year per direction in the top/bottom deciles; the editorial decision stays open for the Chalkbeat checkpoint per D3. For trend, the D1 recommendation is D4(iii) on the winning measure: direction + magnitude of the change in within-year (c) z across 2022-23 → 2024-25 (three scored points, no slope p-values), with the 2018-19-anchored recovery family (D10) as the descriptive companion (median recovery gap +7.8pp; 16.6% fully recovered at the mean endpoint, 13.1% under dual-endpoint persistence). Cross-method disagreement is the expected finding, not a failure: (a)-vs-(c) top/bottom-decile overlap is ~50% (Jaccard 0.32–0.46), inside the memo's 40–60% band — method choice drives about half of who gets named. The D8 persistence gate yields 125–137 schools per direction for (a)/(c); 20 worse and 18 better schools persist on all five measures (names gated per D9, gated_shortlists.json). The D11 composition-stable screen passes only 3.0% of schools under the frozen range-based drift definition, so it functions as a very tight narrative-priority filter, and D12 synchronized shifts corroborate roughly 30–50% of persistent schools against a high (~31%) chance base rate. Net recommendation to carry into the checkpoint: status = (c) WLS residual-z (race variant to be decided editorially), displayed alongside (a)'s peer-percentile sentence; trend = within-year-z endpoint change plus recovery gap; (d) retired as a status candidate.
2. Results per criterion (D7 order; evaluation.json)
All measures are oriented higher = worse than expectation (a: −peer_pctile; c/d: z). Scored window 2022-23 → 2024-25; 2021-22 is robustness-only (D6 amendment). Universe: D6 district schools, memo-vintage pinned (see Deviations, item 1).
Criterion 1 — Year-over-year stability (evaluation.json criterion1_stability)
Baseline to beat: raw school-level CA levels at r ≈ 0.92–0.94 (A.22). Pearson on each measure's own coverage; the matched raw-CA column is computed on the same schools.
| Measure | 2022-23→2023-24 | 2023-24→2024-25 | raw CA, same schools | 2021-22→2022-23 (robustness) |
|---|---|---|---|---|
| (a) peer pctile, year-specific | 0.837 | 0.858 | 0.927 / 0.936 | 0.783 |
| (a) fixed-vintage 2022-23 | 0.843 | 0.865 | 0.927 / 0.936 | 0.781 |
| (c) without race | 0.826 | 0.840 | 0.922 / 0.932 | 0.745 |
| (c) with race | 0.825 | 0.837 | 0.922 / 0.932 | 0.733 |
| (d) no race | 0.023 | −0.118 | 0.922 / 0.932 | n/a (no 2021-22 scores) |
| (d) with race | 0.017 | −0.108 | 0.922 / 0.932 | n/a |
| raw CA levels, full universe | 0.928 | 0.936 | — | 0.892 |
Read: adjustment costs ~0.08–0.09 of YoY Pearson for (a) and (c) — the known price, quantified. Candidate (d) collapses to ~0: lagged-DV residuals are innovations, serially near-uncorrelated by construction. Per the D7 order, (d) fails criterion 1 as a status measure and is not rescued by later criteria (evaluation.json headline_reads.criterion1_stability).
Criterion 2 — SES-orthogonality vs omitted factors (evaluation.json criterion2_orthogonality)
Correlation with in-model covariates is ~0 by construction and is not evidence; the test is correlation with factors each method never saw. Raw CA is the PEER-diagnostic contrast row. Ranges below span the three scored years.
| Measure | ENI | pct_econ_dis | STH share | log enrollment | borough R² | max abs race-share corr |
|---|---|---|---|---|---|---|
| raw CA (contrast) | 0.59–0.62 | 0.53–0.57 | 0.43–0.53 | −0.40 to −0.41 | 0.044–0.050 | 0.50–0.52 |
| (a) peer pctile | 0.20–0.23 ⊘ | 0.15–0.19 | 0.23–0.32 ⊘ | −0.22 to −0.23 ⊘ | 0.024–0.042 ⊘ | 0.18 |
| (c) without race | −0.002–0.018 | 0.003–0.015 | −0.015–0.017 | −0.018–0.005 | 0.017–0.033 ⊘ | 0.08–0.13 ⊘ (race omitted) |
| (c) with race | −0.004–0.014 | −0.004–0.009 | −0.010–0.015 | −0.006–0.011 | 0.037–0.051 ⊘ | ≤0.02 |
| (d) no race | abs ≤ 0.012 | abs ≤ 0.011 | abs ≤ 0.012 | abs ≤ 0.016 | 0.003–0.026 ⊘ | ≤0.07 ⊘ (race omitted) |
| (d) with race | abs ≤ 0.014 | abs ≤ 0.014 | abs ≤ 0.013 | abs ≤ 0.017 | 0.002–0.021 ⊘ | ≤0.03 |
⊘ = a genuinely omitted factor for that measure (evaluation.json criterion2_orthogonality.test_status): for (a), ENI/STH/size/borough are never in the K-NN distance; for (c)/(d) without race, race shares and borough are omitted; with race, only borough is omitted. Per A1: borough is now genuinely held out (its R² 0.017–0.033 is real evidence, not by-construction), and grade mix (gmx, spec F) is in-model — corr(z, grade-mix-expected) ≈ 0.01–0.03, down from 0.11–0.13.
Read: (c) passes everywhere, including on its omitted factors — the without-race variant leaves max |race-share corr| at 0.08–0.13 (vs 0.50–0.52 raw), borough R² 0.017–0.033 (a genuine held-out check now that borough is dropped — small, so the measure isn't geographic), and grade-mix correlation ≈ 0 (controlled via gmx, A1). (a) retains sizable omitted-factor structure (STH 0.23–0.32, ENI 0.20–0.23, size −0.22): K-NN matches 7 demographic shares but never saw STH, ENI, or enrollment. This is the criterion that separates (a) from (c).
Criterion 3 — Cross-method agreement (evaluation.json criterion3_agreement)
REL-Midwest-style measure-comparison framing (external citation still [unverified-in-repo] per the design memo). Deciles recomputed within each pair's common coverage so list sizes are equal (k = 139–142). Memo expectation: ~40–60% list overlap is the finding, not a failure. Ranges across the three scored years:
| Pair | rank corr (Spearman) | worse-decile Jaccard | better-decile Jaccard |
|---|---|---|---|
| (a) vs (c) without race | 0.77–0.78 | 0.33–0.37 | 0.35–0.42 |
| (a) vs (c) with race | 0.76–0.79 | 0.32–0.35 | 0.34–0.46 |
| (c) without vs with race | 0.94–0.96 | 0.66–0.71 | 0.67–0.79 |
| (a) vs (d) variants | 0.40–0.54 | 0.19–0.25 | 0.10–0.20 |
| (c) vs (d) variants | 0.48–0.63 | 0.20–0.31 | 0.16–0.29 |
| (d) no race vs with race | 0.988–0.995 | 0.87–0.92 | 0.86–0.90 |
Read: (a)-vs-(c) decile overlap (overlap_n/k, e.g. 70/139 worse in 2022-23) runs ~50%, inside the memo's expected band — the method choice itself drives roughly half of who gets named. (d) agrees with nothing outside its own family, consistent with its criterion-1 failure.
Criterion 4 — NYC-official cavg benchmark (evaluation.json criterion4_cavg_benchmark)
Data surprise (degeneracy): the ingested cavg_student_chronic_absent takes only 1–2 distinct values per report-year × band — SY 2022-23: EMS {34, 30} (890/352 schools), HS {36} (445 schools); SY 2021-22: EMS {37, 34}, HS {42}. It behaves as a citywide level-group constant, not a school-specific comparison-group average, so school-level Pearson (0.17–0.28 across candidates) is a degenerate point-biserial test and cannot adjudicate between candidates. Criterion 4 was downgraded to a level-calibration check (logged as a deviation; §7 item 8).
Band-matched comparison, single-band schools only (80 multi-band schools excluded), SY 2022-23 scored (n = 1,194–1,233 per row); EM band reconstruction excludes PK (validated: 64% of schools within 0.5pp of the snapshot's own rounded student value vs 34% including PK):
| Expectation (SY 2022-23) | n | Pearson vs cavg | median abs diff (pp) | mean diff exp − cavg (pp) |
|---|---|---|---|---|
| (a) peer-group mean | 1,233 | 0.276 | 8.90 | +5.24 |
| (c) without race, fitted | 1,194 | 0.183 | 8.96 | +5.71 |
| (c) with race, fitted | 1,194 | 0.166 | 9.76 | +5.77 |
| (d) no race, fitted | 1,194 | 0.186 | 10.37 | +6.10 |
| (d) with race, fitted | 1,194 | 0.186 | 10.49 | +6.07 |
| contrast: school's own band rate | 1,250 | 0.173 | 10.44 | +4.42 |
Read: all candidates' mean expectations sit 2–7pp above the cavg constants — but so does the schools' own band-rate mean (+4.4pp), so the offset lives in the benchmark's basis/universe, not in our models (group_calibration rows confirm this per level group). Within the scored window this criterion judged a single year (2022-23), as pre-committed in §5 of the design memo; 2021-22 robustness rows show the same pattern (Pearson 0.25–0.38, offsets +5 to +6pp). Update 2026-06-12: criterion 4 is revived in §9 — cavg turns out to be NYC's "City:" slot (degenerate by construction, our varname-selection artifact); the real school-specific "Comparison Group*" benchmark is pavg_chronic_absent_* (SY 2023-24/2024-25), recomputed there.
Criterion 5 — Interpretability (qualitative; tie-break only)
Verbatim from evaluation.json criterion5_interpretability_qualitative: (a) wins the one-sentence test ("Among the 40 schools most demographically similar to X, it had the Nth highest chronic absenteeism") but its K/weights/Jaccard choices are invisible (the K=20/80 sensitivity is now resolved — Deviations item 7, k_sensitivity.json: typical shift ~5pp and no K rescues (a)). (c) needs one clause more and a "statistical model" footnote but stays honest. (d) is genuinely hard to state as a status claim — "did better/worse than its own track record predicts, given composition" — and with a lag coefficient of ~0.69–0.85 (candidate_d_meta.json) it heavily anchors on last year. Criterion 5 is allowed to break ties only; criterion 2 already separated (a) from (c), so it does not decide the winner here. It does motivate keeping (a) as the display companion.
3. Recovery family (D10; recovery_meta.json, recovery.csv)
Universe 1,463 district schools; primary gap defined for 1,448. Recovery gap (mean-of-2023-24/2024-25 endpoint − 2018-19): p10 −4.2pp · median +7.8pp · p90 +18.0pp. Fully recovered (gap ≤ 0): 16.6% at the mean endpoint, 13.1% under the dual-endpoint persistence rule — the honest headline (both shares over the 1,448 schools with a defined mean-endpoint gap, not the 1,463 universe). Fraction metric defined for 1,326 schools (surge ≥ 5pp and an endpoint defined; 1,337 schools have surge ≥ 5pp unconditionally). A.11 flag applies to 1,121 schools (any grades 1–8 + a 2024-25 endpoint); the 2023-24-only endpoint variant is computed throughout. Data wrinkle (logged): 246 schools have their CA peak outside the pinned 2020-21→2022-23 window (139 peak in 2023-24, 107 in 2024-25) — for them "recovery fraction" is ill-defined and the gap is still rising; they are flagged in recovery.csv, not silently included.
4. Qualifier screens (D11/D12; qualifiers_report.txt, qualifiers.csv)
D11 composition-stable: 44 of 1,450 evaluable (3.0%) under the frozen range-based drift definition — far tighter than anticipated; binding variables are poverty (321 schools), ELL (298), ENI (206). Sensitivity logged, not tuned: the memo did not pin range-vs-adjacent-year drift; max adjacent-year drift passes 27.0%. Recommend an explicit D2 amendment choice at review rather than silent reinterpretation. D12 cohort coherence: ~760 schools scored per year-pair (≥4 transitions, N≥30); coherence ≥0.8 for 44–55%; synchronized shifts 43–55% per pair (improvement/worsening ≈ even), 642 schools in any pair — a high base rate (~31% chance baseline), so coherence corroborates (30–50% of persistent outliers) but cannot identify on its own. Citywide per-transition medians (the U-shape maturation netting) are tabulated in the report file.
5. Flip-diffs (flipdiffs.csv, 8,420 rows)
Method choice is material: (a) vs (c) flips ~130–145 schools per year per direction (≈ half of each ~141-school decile list); (c) with-vs-without race flips ~25 per year per direction; (d) variants agree with each other (0.99/0.89) but with nothing else. All flips carry N, proclivity, grade band, composition-stable and coherence columns for the checkpoint review.
6. Robustness
(c) with/without race: z rank correlation 0.944–0.973 per year — the editorial choice moves tails, not the measure. 2021-22 inclusion (≥2-of-4) is monotone: adds schools only (counts in evaluation.json d8_gate). Fixed-vintage K-NN trends tracked the year-specific variant closely (criterion 3 within-family). (c) fit quality (F): weighted R² 0.57–0.61 per year (without-race 0.57–0.59, with-race 0.60–0.61) — notably above the descriptives memo's 0.46 (richer covariates: ENI, STH share, grade-mix index, admission terms; weighting). K-sensitivity (k_sensitivity.json, 2026-06-17): varying K to 20/80 moves (a)'s peer percentile a median ~3.75–5.0pp (p90 ~11–15pp), and (a)'s omitted-factor correlations are monotone in K — STH 0.24→0.28→0.34 across K=20/40/80 (2024-25) — so no K clears the criterion-2 |r| ≤ 0.03 bar (Deviations item 7, resolved).
7. Deviations log (D2; collated)
- Analysis-table vintage pinned at build time (see
analysis_table_meta.json). - Topic-tag Jaccard included where tags exist; per-year tags unavailable → single tag set reused (peer-group meta).
- D11 "max |Δ|" operationalized as range over observed years; adjacent-year alternative changes pass rate 3.0% → 27.0% (amendment needed).
- D12 churn-flagged transitions included (memo says flagged, not excluded); sign ties count toward neither.
- Recovery peak window pinned 2020-21→2022-23; 246 schools peak later (flagged).
- cavg benchmark degenerate (1–2 distinct values per report-year × band) — criterion 4 downgraded to level-calibration; this is a data finding about the Snapshot feed, ledger candidate below.
- RESOLVED (2026-06-17): K=20/80 K-NN sensitivity reproduced (
k_sensitivity.py→k_sensitivity.json; the K=40 rebuild reproduces the stored peer groups 1442/1442 every year — the correctness gate). Typical per-school shift K40↔K20/K80 is median ~3.75–5.0pp, mean ~5–6pp, p90 ~11–15pp — about half the design memo's team-recalled "~10–15pp," which is the upper tail of movers, not the central tendency. (a)'s omitted-factor correlations grow monotonically with K (2024-25: STH 0.24→0.28→0.34, ENI 0.14→0.23→0.31, log-size −0.18→−0.23→−0.29 across K=20/40/80) and stay an order of magnitude above the criterion-2 |r| ≤ 0.03 bar at every K — no K rescues (a), so the "(a) display-only" call is robust to K. - 2026-06-12: criterion 4 recomputed against the real
pavg_*benchmark (§9); required one SELECT-only DB extract (grades_for_pavg.csv), a deviation from the header's no-DB statement (which still holds for §1–8).
8. Proposed ledger entries (all pending review)
- A.23 (preliminary): bake-off outcome — (c) WLS residual-z selected on pre-committed criteria; (a) fails omitted-factor orthogonality as a standalone status measure (STH 0.23–0.32, ENI 0.20–0.23, size −0.22); (d) not a status measure (YoY ≈ 0).
- A.24 (verified-candidate): the Snapshot's
cavg_student_chronic_absentis degenerate as ingested — 1–2 distinct values per report-year × band (e.g. 2023 EMS {34, 30}, HS {36}) — i.e. a level-group constant, not a per-school 40-peer mean; NYC-official benchmark unusable for school-level adjudication in these years. Superseded 2026-06-12 by the split A.24a/A.24b proposal in §9 (cavg = "City:" slot by construction; school-specificpavg_*benchmark exists for SY 2023-24/2024-25 and was scored there). - A.25 (preliminary): recovery distribution — median school +7.8pp above pre-pandemic; 13.1% fully recovered under dual-endpoint persistence.
- A.26 (preliminary): D11 operationalization sensitivity (3.0% vs 27.0%).
9. Criterion 4 revived — NYC's model-based Comparison Group benchmark (pavg, SY 2023-24/2024-25)
Added 2026-06-12. Inputs: criterion4_pavg.json (+ revive_criterion4.py, grades_for_pavg.csv); scout findings in ../compgroup-scout/SCOUT_NOTES.md. Deviation from the header note: this section used SELECT-only DB access (one per-grade CA extract, grades_for_pavg.csv); §1–8 remain DB-free as stated.
The reframe (per SCOUT_NOTES). The §2 criterion-4 benchmark was the wrong varname for the construct it wanted — a measurement artifact ours, not NYC's. NYC's layout TSVs + Snapshot bundle show cavg_* feeds the "City:" display slot (city_value/desc_value_city): a city average by school type, degenerate across schools by construction, exactly as §2 observed. The "Comparison Group*" slot (comp_value/desc_value_comp) is fed by pavg_* — for chronic absence, pavg_chronic_absent_{ems,hs}_all, published for report years 2024/2025 only (SY 2023-24/2024-25) and already ingested in the repo's data/quality/snapshot-values.json. Membership lists for criterion 4's strongest (overlap) form never existed as an artifact: pre-2023 the Comparison Group was a student-matching construct (50 matched students per student, pooled — NYC's "Data Explained: Comparison Groups" page, the explainer the Snapshot footnote links to), and from report year 2023 it is a regression-based counterfactual. NYC's School Quality Reports Educator Guide states it "worked with MIT Blueprint Labs to develop an updated methodology for Comparison Groups beginning in the 2023 School Quality Reports," in which "student outcomes are regressed on enrolled school indicators," controlling for "student demographics, baseline student achievement, and grade fixed effects" (verbatim, 2023-24 EMS guide PDF, Comparison Group methodology section; the change first appears in the 2022-23 guide). The reader-facing framing — outcomes "if students at a given school had instead enrolled at a random school in the NYC Public School system" — and the confirmation that "for years before 2023, the comparisons are based on a student-matching method" are on NYC's dashboard popover (SCOUT_NOTES §§1, 4–5; raw dash_info_comparison_group.html). Either way there is no school list and no group n. Scope note: the Guide details this regression for Student Achievement metrics; chronic absenteeism is a Supportive-Environment metric, so the Guide establishes the methodology change itself, while the chronic-absence Comparison Group value we benchmark against (pavg_*) carries the same model framing per the dashboard. This resolves the design memo's [unverified-in-repo] flag on the SQR Educator Guide / Blueprint-Labs methodology change — verified, dated to report year 2023, sourced to NYC's own Educator Guide (update the flag in 03_bakeoff_design.md at de-draft time). So criterion 4 is revivable only in benchmark-correlation form, for the two report years where pavg exists — which sit inside the scored window (2023-24, 2024-25), unlike §2's 2022-23-only cavg run.
Decode and orientation confirmation. pavg shares the A.7 complement convention (feed value = round(100 − grade-band CA rate); EMS = K-8 excl PK, HS = 9-12); decoded benchmark CA = 100 − pavg. Confirmed against the schools' own val_chronic_absent_* vs DB band-reconstructed rates: the decoded value lands within 1.5pp of the actual band rate for 96.8–98.2% of schools (per year × band, n = 380–1,007), while the raw value does so for only 1.1–1.8% — orientation is unambiguous. Examples (SY 2023-24): 01M020 feed val 44 → decoded 56.0 vs actual band rate 56.0; 01M448 feed val 67 → decoded 33.0 vs 33.0; 01M015 feed val 48 → decoded 52.0 vs 51.7.
School-specificity (vs the §2 constants). Decoded pavg on the district universe, CA scale:
| SY × band | n | distinct values | sd | p10 / p50 / p90 |
|---|---|---|---|---|
| 2023-24 EMS | 1,007 | 47 | 11.7 | 11 / 30 / 41 |
| 2023-24 HS | 388 | 45 | 8.3 | 30 / 43 / 51 |
| 2024-25 EMS | 1,001 | 48 | 7.8 | 22 / 35 / 41 |
| 2024-25 HS | 380 | 47 | 9.6 | 27 / 42 / 51 |
This is a genuinely school-varying benchmark (vs cavg's 1–2 distinct values per year × band), so school-level correlation is now a meaningful test.
Basis note. pavg is grade-band-based (K-8 / 9-12); our candidates are school-wide. The headline therefore compares school-wide expectations only for single-band district schools (CA-reporting grades entirely within K-8 or within 9-12 — the large majority; 81 multi-band schools excluded in 2023-24, 87 in 2024-25). Candidate (a) expectation = unweighted mean of stored K-NN peers' school-wide CA (self excluded, ≥5 peers); (c)/(d) = WLS fitted CA. pavg is integer-rounded in the feed (±0.5pp quantization floor on diffs).
Headline table — candidate expectations vs NYC's decoded pavg expectation (single-band district schools; Pearson r pooled and per band; diffs in pp, expectation − pavg):
| Expectation (SY 2023-24) | n | r pooled | r EMS | r HS | Spearman | median |diff| | mean diff |
|---|---|---|---|---|---|---|---|
| (a) peer-group mean | 1,227 | 0.631 | 0.551 | 0.746 | 0.698 | 5.35 | +5.73 |
| (c) without race, fitted | 1,208 | 0.659 | 0.625 | 0.839 | 0.685 | 5.55 | +5.37 |
| (c) with race, fitted | 1,208 | 0.656 | 0.623 | 0.858 | 0.676 | 5.64 | +5.38 |
| (d) no race, fitted | 1,208 | 0.620 | 0.591 | 0.750 | 0.633 | 8.36 | +5.44 |
| (d) with race, fitted | 1,208 | 0.617 | 0.587 | 0.748 | 0.629 | 8.29 | +5.47 |
| contrast: raw school CA | 1,240 | 0.612 | 0.584 | 0.736 | 0.617 | 8.91 | +5.26 |
| contrast: own band rate | 1,240 | 0.603 | 0.554 | 0.735 | 0.608 | 8.51 | +4.53 |
| Expectation (SY 2024-25) | n | r pooled | r EMS | r HS | Spearman | median |diff| | mean diff |
|---|---|---|---|---|---|---|---|
| (a) peer-group mean | 1,222 | 0.759 | 0.739 | 0.768 | 0.743 | 4.05 | +1.72 |
| (c) without race, fitted | 1,207 | 0.843 | 0.864 | 0.871 | 0.838 | 3.80 | +1.10 |
| (c) with race, fitted | 1,207 | 0.849 | 0.871 | 0.897 | 0.838 | 3.87 | +1.12 |
| (d) no race, fitted | 1,207 | 0.781 | 0.807 | 0.760 | 0.766 | 5.33 | +1.15 |
| (d) with race, fitted | 1,207 | 0.784 | 0.809 | 0.766 | 0.768 | 5.23 | +1.16 |
| contrast: raw school CA | 1,226 | 0.790 | 0.813 | 0.757 | 0.775 | 5.93 | +1.22 |
| contrast: own band rate | 1,226 | 0.798 | 0.815 | 0.757 | 0.784 | 5.80 | +0.57 |
Benchmark calibration context. Decoded pavg correlates with the school's own band rate at 0.603 (2023-24) and 0.798 (2024-25), with mean offsets −4.53pp and −0.57pp (pavg below own rate) — NYC's model expectation is itself anchored to school composition/outcome levels, and 2024-25 is the better-calibrated year. Data wrinkle (logged, cause not established): 13.4% of ry2024 EMS pavg values decode to CA < 15 (minimum 0 — an expectation of zero chronic absence; pavg = 100 in the feed); the ry2025 share is 1.7%. (Denominator: 159/1,184 EMS-pavg schools feed-wide — the basis that yields 13.4%/1.7%; on the §9 district single-band set the ry2024 share is 10.4%.) Sensitivity excluding the ~92 tail schools from the 2023-24 single-band set: (a) 0.631 → 0.657, (c) without race 0.659 → 0.619, raw 0.612 → 0.573 — the tail does not change the 2024-25 story and makes 2023-24 mildly favor (a) instead of (c); i.e., 2023-24 cannot adjudicate between (a) and (c) (gap ±0.03–0.05, direction flips with the tail), while 2024-25 separates them clearly.
Read / verdict. Against NYC's own published school-specific expectation, candidate (c)'s fitted values agree best overall: highest Pearson in both years pooled (0.656–0.659 vs (a)'s 0.631 in 2023-24; 0.843–0.849 vs 0.759 in 2024-25) and the smallest median absolute gap (3.8–5.6pp vs (a)'s 4.1–5.4 and (d)/raw's 5.2–8.9). In the cleaner benchmark year (2024-25), (c) beats not only the other candidates but the school's own rate — NYC's model and our WLS model converge on what they expect of a school beyond what the school's current outcome alone implies. (a) is competitive (and edges (c) on Spearman and in the tail-excluded 2023-24 view) but never separates above it. This corroborates the §1–2 selection of (c) under the design memo's criterion 4 ("does our peer expectation track NYC's published comparison-group average; large divergence = investigate"): no large divergence, and the ranking is consistent with criteria 1–2. Caveats: two years only, integer-rounded benchmark, single-band restriction, and the ry2024 EMS tail; this section is corroborative, not selection-driving (the D7 order already decided on criteria 1–2).
Criterion 4b — directional agreement (added 2026-06-17; comp_group_direction.py → comp_group_direction.json). §9 above correlated expected levels (our fitted vs decoded pavg); this asks the sharper question: when our metric calls a school above or below expectation, does NYC's comparison group put it on the same side? Per school we compare two gaps — our gap = school-wide CA − (c) fitted (the residual); NYC gap = the school's chronic-absenteeism rate minus its comparison group, both on NYC's grade-band basis (= [100 − Snapshot val] − [100 − pavg], so it reconciles with the Snapshot page; the school side is band/PK-excluded to match pavg, not school-wide); worse = gap > 0. Single-band district schools, candidate (c) without race:
| SY | n | overall directional agreement | resid ↔ NYC-gap (Pearson) | flagged-worse (z≥2) also worse-vs-comp | flagged-better (z≤−2) also better-vs-comp |
|---|---|---|---|---|---|
| 2023-24 | 1,208 | 71.9% | 0.64 | 22/22 (100%) | 38/42 (90%) |
| 2024-25 | 1,207 | 79.9% | 0.82 | 26/26 (100%) | 34/34 (100%) |
Read: at the outlier level the two essentially never disagree — every worse-than-expected school our metric flags also sits above its NYC comparison group (both years; median gap ~18–20pp), and flagged better-outliers 90–100%. Overall agreement is lower only because disagreements concentrate among non-flagged schools sitting near both expectations, where a small calibration difference flips the sign. The 2023-24 asymmetry (more "we say better / NYC says worse" — 232 vs 72 in the 2×2) tracks our own §9 calibration finding that decoded pavg runs −4.53pp below schools' own band rate in 2023-24 (vs −0.57pp in 2024-25; the 13.4% pavg<15% tail) — NYC's 2023-24 bar sits low, mechanically pushing more schools onto its "worse" side. This is our observation; NYC documents no such adjustment and the cause is not established. Net: NYC's comparison group is a strong external corroboration of named outliers for 2023-24/2024-25 — the validation role criterion 4 was meant to play, now confirmed directionally, not just at the expectation level.
A double-confirmed investigate set follows — schools extreme on our z (|z|≥2) and a meaningful comp gap (|CA − pavg| ≥ 10pp) in the same direction: 22 worse / 32 better (2023-24) and 26 worse / 25 better (2024-25). The reverse cut (rank by comp gap, then read z) is a weaker signal — the raw pp gap is not size- or spread-aware the way z is — so it carries enrollment, a small-N flag, and a within-year/band standardized gap; in the cleaner 2024-25, 17 of the 25 largest-gap schools are also our flagged outliers (8 of 25 in the noisier 2023-24). The most robust subset — double-confirmed in both comp years (hence D8-persistent and NYC-corroborated twice over) — is 9 worse / 18 better schools. Named lists are gated (D9): per-school nyc_comp {gap, agrees} now annotates gated_shortlists.json, and the investigate sets (forward, reverse, and the persistent-both-years subset) live in comp_group_outliers.json.
Revised A.24 ledger proposal (replaces the §8 A.24 bullet):
- A.24a (verified): Varname semantics — Snapshot
cavg_student_chronic_absentfeeds the "City:" display slot (city average by school type; constant across same-type schools to 14 digits in dashboardcval, 2019–2025), and was never labeled "Comparison Group" in the UI; the "Comparison Group*" slot ispavg_*, school-specific, published for chronic absence in report years 2024/2025 only. No comparison-group membership list exists as a fetchable artifact in any year — pre-2023 the construct was student-level matching, post-2023 a model-based counterfactual with no group — so building our own K-NN peer groups was not a duplication of an available official product; there was no official membership product to prefer. (Evidence: layout TSVs + bundle render strings + endpoint sweep, SCOUT_NOTES §§2–5; the §2 degeneracy finding stands, reframed as by-construction.) - A.24b (preliminary): Benchmark correlations — against decoded
pavg(= 100 − value, A.7 complement convention; orientation re-confirmed at 96.8–98.2% within 1.5pp), single-band district schools: (c) fitted r = 0.65/0.84 and median |diff| 5.5/3.6pp across SY 2023-24/2024-25 (n = 1,208/1,207), vs (a) peer mean 0.63/0.76 and 5.4/4.1pp; all candidates sit within ~5pp of NYC's expectation on average, with the offset concentrated in the ry2024 EMS low-CA tail (13.4% of values decode to CA < 15; cause not established). NYC's model expectation agrees more with (c) than (a) in the year that can adjudicate.