Source document

docs/writing-style.md

Served verbatim from the project repository. Internal working document conventions apply: documents may reference file paths, branch names, and findings-ledger anchors from the repo.

Writing style guide

This guide governs the tone of analytical and editorial writing across the project — story stubs, "Claude's Answer" investigations, methodology notes, explainers, and any other prose that interprets data for a reader.

The target voice sits between 538, The Economist, and an empirical education-policy blog: numerate, dispassionate, surprised-by-data rather than animated-by-conviction.

Core principle

Describe; do not prescribe. Report what the data shows, name the candidate explanations, surface the trade-offs, and let the reader draw the value judgment. We are not the policy desk. We are not the editorial board.

When in doubt, swap "X is bad" for "X is uncommon / large / outside the expected range," and let the reader supply the verdict.

Things to avoid

1. Assuming a political or ideological frame

Don't treat any of the following as self-evidently good or bad. These are contested:

  • Diversity / integration / segregation
  • Test prep, tutoring, "preparing for the G&T test"
  • Charter schools vs district schools
  • Selective admissions, screening, gifted programs
  • Accommodations for ELL / SWD students
  • Opt-out movements, standardized testing
  • Public-employee unions
  • Funding levels, "fair" funding
  • The role of parents vs schools in attendance
  • Whose growth matters most. Helping a student go from below-grade to proficient is one form of value-add; helping a strong student reach an advanced level is another. The story should not write as if proficiency-rate progress is the only kind that counts, or that wealthy-area schools "have nowhere to grow" because they're already at the proficient ceiling. Proficient is a low bar; there's plenty of room above it.

A school where one demographic outperforms another is showing a gap, not necessarily a failure. A test that produces low scores for ELL students has a measurable disparate outcome, not a broken instrument. A selective school is selective, not exclusionary. State the fact, attribute the value judgment to its proponents, and stop there.

2. Editorial intensifiers and moralizing

Cut these unless the data genuinely supports the intensity:

  • "complete exclusion," "near-unanimous rejection," "stratospheric"
  • "legitimately important," "unsung," "the real story"
  • "stark," "shocking," "alarming," "troubling," "disturbing"
  • "should," "must," "needs to," "ought to" (when applied to policy)
  • "failure to serve," "leaving behind," "left behind"
  • "powerful equity story," "the equity question"
  • "the metric needs supplementation" (the metric measures what it measures)

Prefer: "large," "uncommon," "outside the typical range," "an order of magnitude smaller than X."

3. Loaded characterizations of people and choices

  • Don't describe families "preparing for the G&T test" as if that's a shortcut or unfair. Test prep is a choice some families make.
  • Don't describe wealthy parents "pulling kids" for travel as irresponsible. It's an observation about attendance behavior.
  • Don't describe an opt-out school as either heroic or negligent. Note the pattern and the consequence for the metric.
  • Don't describe a charter "pushing teachers out" without evidence — say "tenure averages 3-4 years."

4. Causal language that outruns the data

Most of our findings are correlational. Use:

  • "is associated with," "tracks with," "co-occurs with"
  • "consistent with," "would be predicted by"
  • "may reflect," "could be," "is one possible explanation"

Avoid: "drives," "causes," "produces," "results in" — unless you have identification (RCT, IV, RD, etc.).

5. Policy recommendations

Don't write "NYC DOE should..." or "policy should..." Instead:

  • Surface the trade-off ("$X buys Y, at the cost of Z")
  • Note who already advocates for each side
  • Let the reader weigh it

If a fact has policy implications, name them as implications, not as mandates. "If a reader values A, the data supports B. If a reader values C, the data supports D." Both framings should appear when both are defensible.

Things to keep

  • Concrete numbers. Always cite the school, the year, the n.
  • Named caveats. "Selection bias is the obvious alternative" is good. "Selection bias" by itself is weaker.
  • Falsifiable claims. "X correlates with Y at r=0.4" is better than "X seems related to Y."
  • Honest uncertainty. "We can't isolate this from the test recalibration" is better than glossing over it.
  • Surprise. If a finding contradicts the consensus, say so plainly — but describe the contradiction, don't editorialize about who was wrong.
  • Reporting directions for journalists. "Worth visiting X" or "Compare Y and Z" is fine. "Story angle: ..." is fine.

Worked examples

Before / after: ELL at specialized HS

Before (too moralistic):

ELL exclusion is total. Across ALL 8 specialized schools combined, ELL enrollment is statistically zero. ... This is not a small or borderline pattern. It's complete exclusion. An ELL student in NYC has essentially zero path to a specialized HS even if their math/quantitative ability is exceptional, because the SHSAT is administered in English with limited accommodations.

The SHSAT-reform conversation has focused on Black and Hispanic under-representation (legitimately important). ELL exclusion is parallel, unmentioned, and more complete.

After (neutral):

ELL enrollment at specialized HS is near zero. Stuyvesant, Bronx Science, Brooklyn Tech, Staten Island Tech, and Brooklyn Latin all report 0.0% ELL; the four smaller specialized schools 0.0-0.3%. NYC's overall HS ELL share is ~14% (~35,000 students). The mechanism is direct: the SHSAT is administered in English, and most ELL students who could clear the cut-score in their native language can't clear it in English.

The SHSAT-reform debate has centered on Black and Hispanic share. ELL share is a separate, larger gap from citywide composition, and has drawn less attention. Whether that's the right priority depends on what reform is trying to achieve.

Before / after: Black-white within-school gap

Before:

The school's reputation rests on the kids whose home environments would carry them through any school. ... failure to serve Black students.

After:

The school's aggregate score is anchored by the subgroup that scores highest. The within-school Black-white gap of 77 points means the Black-student subgroup is scoring well below the citywide Black-student average, even though the school appears strong on its overall numbers. Whether that reflects within-school tracking, prior preparation differences, or instructional choices is what a reporter would need to disentangle.

Before / after: Anderson School

Before:

The school has approximately zero outcome variance from its inputs. For learning about what good teaching looks like, Anderson is uninformative. For understanding inequity in access, it's central — NYC has very few K-8 G&T programs, Anderson is the most-coveted, and admission is dominated by families that prepared for the G&T test.

After:

Anderson's outcome is closely tracked by its admissions screen — students enter scoring at the 90+ percentile of the G&T assessment, and the school's outcome distribution is at the test ceiling. That makes Anderson useful as a case study in selection effects and less useful as an example of instructional practices that generalize. Admission depends on the kindergarten G&T test, and families with the resources to prepare for that test are over-represented.

Before / after: Wealthy schools with absenteeism

Before:

These should not have absenteeism over 50%. Something has changed in upper-middle-class attendance culture.

After:

These schools serve mostly advantaged populations and yet post >50% chronic absenteeism — well above what their demographics alone would predict. The pattern is consistent with a shift in family attendance norms post-COVID: more elective absences for travel, mental-health days, and minor illness. It is one of several possible explanations.

Quick checklist before publishing

  • Did I use "should" or "needs to" about policy? Cut or attribute.
  • Did I call a pattern "complete," "total," "stark," or "shocking"? Replace with a number.
  • Did I treat diversity, test prep, or accommodations as good or bad on their face? Reframe as a trade-off.
  • Did I assert causation? Downgrade to correlation unless I have identification.
  • Did I name the alternative explanation, even the one I find less compelling?
  • If a left-leaning reader and a right-leaning reader both read this, would each feel the data was reported fairly? If not, edit.

Why this matters for this project

This site is an analytical resource for parents, journalists, and researchers across the political spectrum. Readers who feel the framing is partisan will discount the findings — even findings that have nothing to do with politics. Neutrality is not a stylistic preference; it's how the data stays useful to everyone.