Validation methodology — Philadelphia
Per-metric definitions, validation results, and known gaps. Mirrors
verify/METHODOLOGY.md (NYC) in shape.
The four checks (terms used throughout)
Identical to KX's NYC framing:
- (i) Completeness — are years or schools missing? Surfaced as "source-only" / "served-only" cells.
- (ii) Correctness — base case — does the DB match an independent re-read of the exact source file the loader ingested? (the binding correctness gate)
- (iii) Correctness — spot check — does the DB match a public-facing per-school report? Two flavors for Philly:
- (iii-a) DB vs SDP SPREE — same publisher, presentation cross-reference. Deferred until SPREE URL is discovered + ingested.
- (iii-b) DB vs Future Ready PA Index — different publisher, independent oracle. Available day-one for Philly (the gap NYC had to defer).
- (iv) Correctness — computed values — derived metrics (composites, survey rollups). (deferred)
Plus the Acct vs Actual population-cut diff specific to Philly's O.8 decision — see __acct_vs_actual.md.
Validation results — 2026-06-04 snapshot
(i) Completeness — PASS
All 6 currently-validated components show clean coverage. The single
source-only graduation row is a year-encoding edge case (a cohort year that
falls outside our default reconciliation window). See
__completeness.md for the full table.
(ii) Base case — PASS across all 6 components
Per-cell agreement at ±1.0 pp (±0.5 pp for graduation + attendance):
| Metric | Common cells | Matches | Mismatches | Status |
|---|---|---|---|---|
| pssa_all_ela_proficiency (Acct) | 170 | 170 | 0 | PASS |
| pssa_all_math_proficiency (Acct) | 170 | 170 | 0 | PASS |
| pssa_grade3_ela_proficiency (Acct) | 150 | 141* | 0 | PASS |
| keystone_algebra_proficiency (Acct) | 56 | 54* | 0 | PASS |
| attendance_rate_above_90 | 216 | 216 | 0 | PASS |
| graduation_rate_4yr | 81 | 75* | 0 | PASS |
* matches less than common = suppressed cells (correctly excluded from
the match count). No actual disagreements at any cell.
(iii-b) Cross-publisher cross-check — sanity-PASS
SDP PSSA Grades 3-8 (ELA) vs Future Ready Keystone Literature (HS ELA): 169 pairs, median Δ=0.03 pp, 94% within ±1 pp, 99% within ±5 pp.
NB: PSSA grades 3-8 and Keystone Literature HS measure ELA at different grade levels, so they're not expected to agree exactly — only to correlate. The tight agreement is consistent with both being honest ELA measures.
A tighter (iii-b) check would compare PSSA grade-by-grade against Future Ready's per-grade breakouts; the per-grade FRPI columns aren't loaded yet, so deferred to v2.
Acct vs Actual population-cut diff (O.8 gate) — TRIVIAL
Every one of 4,535 paired cells shows Δ = 0. At the school × subgroup rollup grain for 2024-25, SDP's Acct (accountability cohort) and Actual (all tested) cuts are byte-identical.
This is itself a substantive finding. The conceptual cohort-vs-tested
distinction that drove NYC's reconciliation gap (per
verify/METHODOLOGY.md — 67–73% snapshot agreement with wide tails) simply
doesn't move the published number in Philly for 2024-25. Two interpretations
to test in future years:
- SDP's pipeline produces both files from the same intermediate — the "two cuts" naming is a vestige of a possibly-different reporting intent that never produced numerically distinct outputs at the school-aggregate grain.
- The Acct/Actual difference is real at the student level but disappears at school aggregation because the in/out flips cancel.
Either way: for 2024-25 Philly, Acct = Actual at school-level rollup. The harness retains both load paths so that if 2025-26 or later shows a divergence (e.g. SDP changes attribution rules) it surfaces immediately.
Per-metric definitions
PSSA proficiency (all-grades + per-grade)
- Ours (loaded from SDP's PSSA Zip):
% Proficient + Advanced=profadv_scorecolumn, per (school, year, testname, subject, grade, group). Ingested in both Acct and Actual cuts. - (ii) base case: PASS at ±1.0 pp across 320 paired all-grades cells (ELA + Math), 150 grade-3 cells.
- Coverage: SDP district + alternative only (220+ schools). Charter PSSA aggregate values come from Future Ready (Future Ready aggregate column populated).
Keystone (Algebra / Literature / Biology)
- Ours (loaded from SDP's PSSA Zip and Future Ready Performance file):
% Proficient + Advanced. Per O.5, latest-attempt all-students. - (ii) base case: PASS at ±1.0 pp on Algebra (56 paired cells).
- Coverage: HS schools only.
PVAAS growth (Algebra / ELA-Literature / Science-Biology)
- Ours (loaded from Future Ready Performance file): growth-index value per (school, subject, subgroup).
- (ii) base case: not yet wired (deferred — PVAAS data structure makes per-cell reconciliation noisier; first add to v2).
- Coverage: all 298 Philly schools where reported by PDE.
Graduation rate (4-/5-/6-year)
- Ours:
scorecolumn from SDP's grad CSV, indexed by cohort year (9th-grade start year). Future Ready charter-side covers report-year indexed values for the same metric_key (intentional — see loader notes). - (ii) base case: PASS at ±0.5 pp on 4-year SDP cohort 2021-22 (81 paired cells).
Attendance ≥90% rate
- Ours (SDP attendance CSV):
% with 90%+ Attendance (Yearly). - (ii) base case: PASS at ±0.5 pp (216 paired cells, 2024-25).
Persistent attendance rate (Future Ready)
- Ours (Future Ready On Track Measures sheet):
PercentPersistentAttendance(the metric PDE actually publishes; theirPercentChronicAbsenteeismcolumn exists but is NULL statewide in 2024-25). - (ii) base case: not yet wired in
components.yaml(the file structure makes the reconciliation trivial; pending one more iteration).
Out-of-school suspensions (derived: any / multiple / chronic)
- Ours (SDP suspensions CSV → derived per F.10):
oss_pct_any=100 − % zero;oss_pct_multiple=% 2 + % 3 + % 4+;oss_pct_chronic=% 4+(residual). - (ii) base case: not yet wired — these are computed metrics so v1 reconciliation against the raw distribution columns is straightforward but pending.
- (iv) computed-value check: needed once we add it.
PSES survey rollups (teacher + student × 6 domains)
- Ours (PSES topic+subtopic Zip): per-respondent topic-level score derived as mean of subtopic scores per (school, year, topic, respondent).
- (ii) base case: deferred — KX's NYC equivalent took several iterations; mirror that timing.
- Coverage: 207-298 schools per metric.
What's still TBD
- (iii-a) SPREE spot check — needs the SPREE URL discovered + an ingest pass; then a per-metric mapping (
__snapshot.yamlequivalent) and a comparison runner. - (ii) base case for PVAAS (attendance_persistence_rate — the state's Regular Attendance, stored by attendance year — is now covered by
absenteeism_full_qa.py§B; OSS and PSES by__computed_value). - (iv) computed-value checks for the OSS derivations and the PSES rollups.
- A "1-of-N anchor" set of hand-verified schools for each metric (KX's
known_cases:pattern), so a known-good cell is always in the sample.
Authority
This document is authoritative for Philly verification results as of 2026-06-04, with the attendance checks superseded on 2026-09-07 by absenteeism_full_qa.py and the year-aligned cross_publisher_attendance.py (see the corrections log on /philly/methodology/metrics/chronic-absenteeism).
Reports are regenerated by python pipeline_philly/verify/reconcile.py --all and
the companion scripts. Where this diverges from KX's NYC framing, the divergence
is intentional and documented in 05_validation_spec.md.