Source document

docs/agents/pass-log/covid-test-recalibration.md

Served verbatim from the project repository. Internal working document conventions apply: documents may reference file paths, branch names, and findings-ledger anchors from the repo.

Pass log — covid-test-recalibration

Story: data/stories/answers.ts"covid-test-recalibration"


PASS 1 — Editor (2026-05-29)

WHAT I CHANGED

  • shortPreamble led with one specific number (grade-3 math jumped from 53.9% to 61.8%) but didn't anchor the structural finding — that NYS rebuilt the scoring scale in 2022-23 and lifted the cutscore in 2024-25. Restructured: lead with the structural finding, then the illustrative number.

  • TLDR was three sentences but each was a bullet-style claim. Reworked into a paragraph that makes the case in narrative form.

  • "The smoking gun" section heading is too tabloid. Renamed to "Why this is a recalibration, not learning."

  • "What this means for any 'NYC schools recovered' coverage" — this is correct but the section's body reads like a lecture. Tightened.

LEGAL REVIEW NOTES

  • Naming NYSED (NY State Education Department) is safe — they're a state agency, the data is public.
  • The "smoking gun" framing was at risk of implying intentional manipulation; the new framing ("scale was rebuilt") is descriptive and safer.

PASS 2 — Education expert (2026-05-29)

WHAT I VERIFIED OR CORRECTED

  • NYS scale change in 2022-23: confirmed. NYSED released the new "Next Generation Learning Standards"-aligned assessments in 2022-23 on a new scale. The visible drop from ~600 to ~450 in mean scale scores is the scale reset, not a measured ability decline.

  • 2024-25 cutscore adjustment: less publicly documented than the 2022-23 scale change. NYSED's technical reports for 2024-25 confirm proficiency rate changes at the cutscore boundary but the agency hasn't called it a "cutscore adjustment" in public-facing communication. The story should be careful: we're inferring this from the visible signature in the data (flat scale scores + jumping proficiency rates), which is suggestive but not the same as NYSED confirming it.

  • NAEP comparison: NAEP grade-4 math NYC results 2022 → 2024 showed +2 points; grade-4 reading −5. These are different from the reported NYS proficiency rate movements and support the recalibration argument. Verified.

  • Andrew Ho's test-equating framework is the right citation here. His "Vertical Scaling" work at Harvard Graduate School of Education is the foundational methodology.

WHAT I ADDED

  • Caveat that the "2024-25 cutscore adjustment" framing is an inference from the data, not a stated NYSED action — the language should be precise about this.
  • Specific NAEP numbers for 2024 NYC grade-4.
  • Andrew Ho citation strengthened.

PASS 3 — Quantitative (2026-05-29)

WHAT I VERIFIED

  • Scale-score values: the ~600 → ~450 shift is real in the data, visible at every grade tested (3, 5, 8 for both math and ELA). The drop is uniform across grades and subjects, consistent with a scale reset rather than differential measured-ability changes.

  • Mean scale scores 2022-23 → 2024-25 are nearly flat: math grade-5 went 450.3 → 451.9 → 454.6 (+4.3 over two years). ELA grade-5 went 448.5 → 444.0 → 451.8 (+3.3 net, not monotonic). These small shifts are consistent with normal year-over-year cohort noise, NOT with the 5-15 point proficiency rate jumps in the same data.

  • The proficiency vs scale score divergence is the strongest available evidence of a cutscore shift. Verified at every grade-subject combination.

WHAT I CORRECTED OR QUALIFIED

  • The "two recalibration events, not one" framing is statistically defensible but worth qualifying: events 1 (scale reset) and 2 (cutscore movement) are different in kind, and the story should make that distinction clearer. Reset = new scale construction, publicly announced; cutscore movement = where to draw "proficient" line, less publicly documented.

  • NAEP as external benchmark: the right comparison, but worth flagging that NAEP is sampled (not census), so it carries its own sampling error.

WHAT I FLAGGED FOR DOWNSTREAM

  • Editor (pass 4): I added a "scale reset vs cutscore movement" distinction. Verify it doesn't make the story too technical for the general reader.

PASS 4 — Editor (2026-05-29)

WHAT I CHANGED

  • Adopted the expert's "the 2024-25 cutscore adjustment is an inference, not a stated action" phrasing.
  • Took the quant's "scale reset vs cutscore movement" distinction and folded it into the "two recalibration events" section without letting it run long.
  • Adjusted NAEP numbers per expert's note.

PASS 5 — Education expert (2026-05-29)

WHAT I VERIFIED OR CORRECTED

  • Cross-checked the NAEP 2024 NYC results once more — figures are +2 math, −5 reading at grade 4. (At grade 8, NYC math was flat and reading down 1; these aren't currently in the story but could be added.)
  • One historical note: NYS's previous scale reset (2017-18) is visible in our older data too. Cross-year proficiency comparisons spanning 2017-18 would also need adjustment.

PASS 6 — Quantitative (2026-05-29)

WHAT I VERIFIED

  • All numerical claims still hold.

WHAT I CORRECTED OR QUALIFIED

  • Added a one-line note that NAEP is sampled (the right comparison but carries its own sampling error).

PASS 7 — Editor (final) (2026-05-29)

WHAT I CHANGED

  • Final polish on the line-chart captions (consistent voice).
  • Tightened TLDR by one sentence.

KNOWLEDGE-FILE UPDATE

Added "When a story argues a state agency did X, distinguish between 'stated action' and 'inference from data'" to docs/agents/quantitative-knowledge.md.


FINAL STATUS

Story is publishable. ~1950 words; net change ~+50 words across passes.


PASS 8 — Plain-language (2026-05-29)

WHAT I CHANGED

  • TLDR + shortPreamble fully rewritten. Lead is now "Headlines about New York City schools recovering…" — orients the reader to what this story is responding to before launching into evidence.

  • "Scale score" → "raw test score" everywhere, with inline definition: "raw test scores (called scale scores)" and "the actual underlying measure: scaled numbers from 100 or so on up." The technical name is preserved once parenthetically so the reader knows the term if they encounter it elsewhere.

  • "Cutscore" → "cutoff" / "the score required to be called 'proficient'." Cutscore is education jargon; cutoff is normal English.

  • "ELA" → "English Language Arts" first use, then "reading" for subsequent uses in chart labels and headings. Most readers don't know what ELA stands for.

  • "NYS" / "NYSED" → "New York State" / "the state education agency" on first use; "New York State" thereafter.

  • "NAEP" → "the federal NAEP test" first use, then defined more fully in the body section: "the National Assessment of Educational Progress."

  • "Recalibration" → "test getting easier" / "test change" / "the test being rebuilt" / "the cutoff moving." The word "recalibration" stays in the story ID but doesn't appear in the prose.

  • Chart markers: "Scale reset" → "New test, new scale." More reader-legible.

  • "Vertical scaling," "test equating" kept but with translation ("the technical term is 'test equating'") in whatsNext.

  • "FOIL" spelled out: "file a Freedom of Information Law request."

  • Section headings rewritten:

    • "The 'recovery' over time, by test (grades 3-8)" → "What the 'recovery' looks like year by year"
    • "The smoking gun: scale scores fell by ~150 points overnight" → "Why this is a measurement change, not a learning change"
    • "The flat scale scores vs. jumping proficiency rates" → "Raw scores barely moved — but proficiency rates jumped"
    • "Comparison: year-over-year change rates" → "How big a year-over- year jump is normal?"
    • "Two recalibration events, different in kind" → "Two changes to the test, of two different kinds"
    • "What this means for 'NYC schools recovered' coverage" → "What this means for any 'New York City schools recovered' coverage"
  • Table column headers: "Mean scale score" → "Average raw score." "Δ scale" → "Change in raw score." "Δ %" → "Change in proficiency."

JARGON KILLED

  • scale score (replaced with "raw test score" + inline definition)
  • cutscore (replaced with "cutoff")
  • recalibration (replaced with descriptive phrases)
  • ELA (replaced with English Language Arts / reading)
  • NYSED (replaced with "the state education agency")
  • vertical scaling / test equating (kept but translated)
  • "discontinuity"
  • "the signature of"

ACRONYMS SPELLED OUT (first uses)

  • NYC → New York City
  • NYS → New York State
  • ELA → English Language Arts
  • NAEP → National Assessment of Educational Progress
  • FOIL → Freedom of Information Law
  • GSE (Harvard) → spelled out as "Harvard education researcher"

NEW PLAIN-LANGUAGE PATTERNS I'M ADDING TO THE KNOWLEDGE FILE

  • "Scale score" → "raw test score" with inline parenthetical. The raw-score concept is essential to the story; the jargon name is not.
  • "Cutscore" → "cutoff." Cutscore is industry jargon for "the score required to be called proficient." Replace with normal English.
  • "ELA" → "English Language Arts" first use, then "reading" for most natural-language descriptions. Charts and tables can use either.
  • "NAEP" needs the full name AND a clarifier on first use ("the federal NAEP test, given the same way in every state").

WHAT I LEFT ALONE AND WHY

  • The numerical claims and the structural inference. The plain- language pass is about how it's said, not what's said.
  • Andrew Ho's name — public researcher, appropriate to cite.
  • "Next Generation Learning Standards" — this is the actual program name, not jargon.

Net change: ~1950 → ~2350 words.