NJ School Data
← Data library
Data explainer

NJ school discipline by student group

2024–25 School Performance Report · one school year, no trend

New Jersey publishes, for every district, how many students were suspended, referred to police, and involved in an incident that led to an arrest — broken out by student group and by grade. For the groups a parent most wants to see, the file is blank more often than it is filled. This page is the manual for reading it, and the pattern of the blanks is half of what it has to teach.

1 · The source

Every year New Jersey publishes a School Performance Report for each school and district, and alongside the web pages it posts the whole thing as a spreadsheet. Inside that file are six tabs of discipline and safety records: reported incidents by type, HIB (harassment, intimidation and bullying) investigations by nature, disciplinary removals, police notifications, and arrests. Three of those tabs are broken out by student group and by grade.

The edition here is 2024–25, covering 666 district reporting units. It is one school year. There is no trend on this page and no way to make one, because the file does not carry a comparable prior year.

We show this on district profiles, in the section Which students are disciplined here, joined only by exact zero-padded NJDOE district code — never by name. 576 of 587 rendered districts match. Eleven do not, and they show nothing rather than a guess.

We do not show it on school profiles. The release carries school-level rows for 259 school codes against roughly 2,900 schools on this site, which is not coverage — rendering it would imply a completeness that isn't there.

2 · Three measures, and never one score

A removal is not a police call. A police call is not an arrest. None of the three is a measure of safety.

The district module carries three measures, kept apart and never added:

MeasureWhat NJDOE counts
SuspendedStudents who received one or more suspension of any type — in-school, out-of-school, or both.
Police notifiedStudents involved in at least one incident the school reported to police. A notification is a referral, not a charge and not a finding.
ArrestedStudents involved in at least one incident that led to an arrest.

There is no combined index here, and there will not be one. Adding removals to police notifications to arrests would triple-count the same student and would produce a number — a “school safety score” — that the underlying data cannot support. The three measures also overlap by construction: a student suspended and referred to police appears in both rows.

One quirk you should know about the arrest column: NJDOE's own source fields are named PoliceCount and PolicePercent, while the published layout defines them as all incidents leading to arrest. That mismatch is recorded by the publisher and carried through to our pages as published. We do not reconcile it against the separate police-notification tab, and neither should you.

3 · Five things a cell can be

The most reversible mistake on this data is treating a printed zero as missing, or a withheld cell as zero. They are opposite errors and both change the finding.

Every cell in the file is one of five publisher acts, and we render each of them differently:

What the file saysWhat it meansHow we show it
1.6%A rate NJDOE calculated and printed.1.6%, with the headcount beside it where one was published.
0 · 0.00%An exact zero. The state is saying nobody — not saying nothing.None, with “the state published a zero”.
<5A headcount masked below five — including zero. The rate is withheld, but the bound is real information.Fewer than 5, with “0 to 4 students”.
*Withheld outright. No value, and no bound.Withheld.
n/aA cell the layout does not define. Structural, not suppressed.Not collected.

The two suppression markers do not mean the same thing. Across the whole release, 661,604 suppressed cells carry * and 89,178 carry <5. A tool that collapses both to “no data” throws away a bound the publisher chose to give. We keep the publisher's literal verbatim all the way from the source file to the page.

There is a further wrinkle: <5 only ever lands on a headcount. Its partner percentage in the same row is always a bare *. Read the percentage alone and the bound vanishes on 34,193 district cells. We resolve every cell from the count and the percentage together for exactly this reason.

And <5 includes zero. It reads like “between one and four students”, and it isn't. Grades are mutually exclusive, so a district's published grade headcounts plus its <5 grades have to reconcile to its district total — and in 96 districts they can't. District 0070 suspended one student all year and carries ten grade cells marked <5. At least nine of those ten are exactly zero. Read the marker as evidence that something happened and a small district's grade table inverts completely: ten grades that look like scattered discipline are one child and nine masked zeros. We say “0 to 4 students”.

4 · How holey it is

Suppression runs inversely to group size. The smallest groups — the ones for whom a disparity would be most visible — are the ones most often withheld.

On the surface we render — three measures across twelve student groups in 666 districts, 23,976 cells — this is where the state landed:

Student groupA rateA published zero<5WithheldWithheld %
Non-Binary/Undesignated Gender33061,95998.05
American Indian or Alaska Native39125311,80390.24
Asian, Native Hawaiian, or Pacific Islander34253320192246.15
Two or More Races34057122186643.34
Black or African American59558528153726.88
Female755850638719.37
Male815790638719.37
White7137813661386.91
Economically Disadvantaged7897383661055.26
Hispanic/Latino771798384452.25
Students with Disabilities798800385150.75
All Students1,207791000.00

This is why we never drop a withheld row from a chart. Because the holes are patterned by group size, a chart of only the surviving cells would quietly redraw discipline as a White, Hispanic, Black, economically-disadvantaged and disabled-student phenomenon — those being the five groups whose cells mostly survive. Every group keeps its row, with the reason for its absence written where the value would be.

It also rules out something readers reasonably want. With Black-student cells withheld at 26.88% and White-student cells at 6.91%, any ratio between the two is computed across differently-censored samples. We publish no disproportionality index, risk ratio, or “X times more likely” statement from this file, anywhere. The groups also overlap — a student can be counted as Hispanic, economically disadvantaged and classified as having a disability at once — so the rows never sum to the district row either.

Grade is a different story

By grade the file is nearly complete: between 0.60% and 2.93% withheld. But that completeness is mostly published zeros, not published rates. Grade PK is 78.6% zero and 0.33% rate. The early grades are a substantive finding about who gets suspended, not a data gap — and reading those zeros as missing reverses it exactly.

We show grade on its own rather than crossed with student group. The file does not cross the two dimensions, and a grid of them would be mostly absence rendered in a channel built for magnitude.

Small districts are not censored more — they are recorded as zero

It is tempting to write that small districts get censored. The data says otherwise. Suppression is roughly flat at 29–31% for every enrollment band above 250 students, and the correlation between district size and suppression rate is weak (Pearson r = −0.2756, n = 665). What changes dramatically is the reported rate, which rises about fourteen-fold from the smallest districts to the largest. In districts under 250 students, 97.7% of discipline cells are either a suppression marker or a published zero. The correct claim is that small districts are recorded as zero rather than published as a number.

5 · The withheld headcount

For male and female students, New Jersey hides the number and prints the rate.

This is the single strangest thing in the file, and it decides how the whole surface has to be built. Male and Female each have 12,654 district headcount cells and 12,654 matching percentage cells. NJDOE suppresses 12,467 of the headcounts (98.52%) and only 2,040 of the percentages (16.12%). Every other student group suppresses the two columns at identical rates.

Stranger still: in 8,328 male and 8,506 female cells, the withheld headcount sits next to a percentage of 0.00%. The state starred out a number it had already told you was zero.

So we lead with the rate, not the count. A surface built on headcounts would find sex unreportable in 98.5% of districts, and — worse — would print “withheld” over 16,834 cells where the state plainly said nobody. A rate-led surface keeps sex readable in 83.9% of districts. Where NJDOE published the headcount we show it; where it didn't, we say so in place rather than leaving a gap the reader has to interpret.

One thing the percentage is not: it is not the group's share of the district's suspensions. It is the share of that group's own students, against that group's own enrollment. A district's Black suspension rate and its White suspension rate are two rates over two different denominators, which is what makes them worth reading side by side — and what makes dividing one by the other a different and much shakier operation.

6 · Rates on a handful of students

NJDOE masks 34,193 small district headcounts with <5. But it publishes another 13,939 headcounts of one to four outright, so its own small-cell rule covers only about 71% of the cases. Nine districts publish a student-group rate of exactly 10.00% off a headcount of 1.

On our three-measure surface, 1,960 of the 7,167 published student-group rates — 27.3% — rest on a headcount below five. A single student in a group of ten prints as 10%, which is arithmetic and not a finding.

So we state a floor and enforce it: a rate whose published headcount is one to four keeps its row in the table, with the headcount right beside it, and gets no dot on the chart. The chart declines to draw a mark that invites a comparison the number cannot carry; the table still shows exactly what the state published. Every rate that does carry a dot is shown with its headcount, or with a note that the headcount was withheld.

7 · Caveats that matter

These are reported incidents and the responses adults chose. They measure how a district records and refers at least as much as how its students behave. A district that documents diligently and calls police readily will look worse here than a district that handles the same behaviour informally — and nothing in the file distinguishes the two. This is not a crime rate, not a safety measure, and not a measure of school quality.

8 · What we built

The whole family arrives as an exact, offline, hash-pinned release — this repo downloads and parses nothing. From it we project four district tables, plus a 54-row catalog of what each metric means.

TableRowsGrain
discipline_safety_district_headline8,658District totals — incidents by type, police notifications
discipline_safety_district_hib_nature15,984District × the eight natures of a HIB investigation
discipline_safety_district_all_students25,308District × the publisher's All Students row
discipline_safety_district_group_grade570,342District × student group and district × grade — the table behind this page
discipline_safety_metric_catalog54One row per metric: label, definition, unit, publisher defect

Until 2026-07-30 only the first three existed, and the warehouse had a suppression rate of exactly 0.000% — because those three slices are precisely the ones New Jersey never suppresses. That was a property of what we had chosen to keep, not of the data. The widened table carries all 183,192 district-level suppressions, and value_state and the publisher's literal ride along with every one of them.

The three original tables are asserted byte-identical by the builder itself, so the widening is provably additive.

The issue registry

Every dataset behind this site keeps a structured registry of its known issues — the format breaks, suppression rules, entry errors, and definitional traps we hit while building on it — so that the next person (or agent) who works with this data doesn't rediscover them the hard way. Each issue has a stable id, a machine-readable scope (which years, columns, and tables it touches), and an effect: breaks stops a pipeline, corrupts silently wrongs the numbers, misleads invites a wrong reading of right numbers, context is background you must hold to use the data responsibly. Issues marked ★ are core: read them before any use of this data. The registry is maintained in the ergo format and served in machine-readable form alongside this page — links at the end of this section.

NJ Department of Education, Office of Performance Reports · source confidence A · updated 2026-07-30 · 10 known issues (3 core)

The pitfall: These are separate 2024-25 records of reported incidents and responses, not a combined safety score or evidence of cause, effectiveness, or budget performance.

two-suppression-markers [misleads · mitigated] Suppressed cells carry either `*` or `<5`, and `<5` bounds the value where `*` says nothing at all

Type: suppression — applies to tables: discipline_safety_district_group_grade · columns: reported_literal, value_state · rows: value_state = 'suppressed'

How to spot it: Release-wide the 750,782 suppressed values split 661,604 `*` and 89,178 `<5`. In the district student-group and grade slice the split is 148,999 `*` and 34,193 `<5`. No suppressed cell carries a numeric value.

The misread: Coalescing both markers to NULL or to a single 'suppressed' label. `<5` bounds the headcount below five; `*` states only that the publisher withheld the cell. Collapsing them discards a bound the publisher chose to give. Note the bound runs from zero, not one — see less-than-five-includes-zero.

Full entry, with the story and the numbers, in the served ergo doc.

less-than-five-is-count-only [misleads · mitigated] `<5` never appears on a percentage metric; the percentage partner of a `<5` headcount is always a bare `*`

Type: suppression — applies to tables: discipline_safety_district_group_grade · columns: reported_literal, unit · rows: reported_literal = '<5'

How to spot it: All 89,178 `<5` literals in the release sit on the students-or-incidents unit; the percent unit carries only `*` (322,625 cells). Joining each count metric to its percent partner on the same source row gives the pair ('<5','*') in every student group and every grade.

The misread: Resolving a cell from its percentage alone. The percentage says only `*`, so a reader who never looks at the headcount column throws away the below-five bound on 34,193 district cells.

Full entry, with the story and the numbers, in the served ergo doc.

less-than-five-includes-zero [corrupts · mitigated] The `<5` marker bounds a headcount at zero to four, not one to four — most `<5` cells are masked zeros

Type: suppression — applies to tables: discipline_safety_district_group_grade · columns: reported_literal · rows: reported_literal = '<5'

How to spot it: Grades are mutually exclusive and partition a district, so the published grade headcounts plus the `<5` grade cells must reconcile to the All Students headcount. Among the 574 districts with a published All Students any-suspension count and no `*` grade cell, 96 have an All Students total smaller than their number of `<5` grade cells — district 0070 suspended one student in total and carries ten `<5` grade cells, district 0080 one student and nine. Those cells cannot all contain at least one student, so at least nine of ten are exactly zero.

The misread: Reading `<5` as 'between one and four students', which turns a masked zero into an incident. In a small district that inverts the whole grade table: ten grades each reading 'fewer than 5' suggests discipline spread across every grade when the district suspended one child.

Full entry, with the story and the numbers, in the served ergo doc.

small-cell-rule-is-inconsistent [misleads · mitigated] NJDOE masks 34,193 small district headcounts as `<5` but publishes 13,939 headcounts of one to four outright, so a published rate can rest on a single student

Type: suppression — applies to tables: discipline_safety_district_group_grade · columns: numeric_value · rows: value_state = 'reported' AND unit = 'students-or-incidents' AND numeric_value < 5

How to spot it: In the district student-group and grade slice, 13,939 reported headcounts fall between 1 and 4 while 34,193 comparable cells are masked `<5`; 15,079 reported headcounts are 5 or more. Nine districts publish an American Indian or Two or More Races rate of exactly 10.00% off a headcount of 1. On the three-measure student-group surface, 1,960 of 7,167 published rates (27.3%) rest on a headcount below five.

The misread: Treating a published rate as evidence of a pattern. One student in a group of ten prints as 10%, which is arithmetic and not a finding; and treating the `<5` mask as a guarantee that every small cell is protected.

Full entry, with the story and the numbers, in the served ergo doc.

sex-count-withheld-rate-published [misleads · mitigated] For Male and Female alone, NJDOE suppresses 98.522% of district headcounts and only 16.121% of the matching percentages

Type: suppression — applies to tables: discipline_safety_district_group_grade · columns: value_state, reported_literal · rows: conformed_dimension_label IN ('Male','Female')

How to spot it: Male and Female each carry 12,654 district count cells and 12,654 percent cells. 12,467 count cells are suppressed against 2,040 percent cells. Every other student group suppresses the two columns at identical rates to three decimals. Pairing each count with its percent on the same source row, 8,328 Male and 8,506 Female cells carry a suppressed `*` count beside a published `0.00%` percentage. `spr.discipline.arrests.any-count` is suppressed for 659 of 666 districts for both.

The misread: Building a student-group surface on headcounts. Sex becomes unreportable in 98.5% of districts, and in the 8,328/8,506 zero-percentage cells a count-led reading prints 'withheld' over a cell where the publisher stated an exact zero.

Full entry, with the story and the numbers, in the served ergo doc.

suppression-tracks-group-size [misleads · mitigated] The smallest student groups are withheld most: 98.522% of Non-Binary cells and 92.327% of American Indian cells against 16.603% for Students with Disabilities

Type: coverage — applies to tables: discipline_safety_district_group_grade · columns: value_state · rows: dimension_kind = 'student-group'

How to spot it: Each student group carries 25,308 district cells. Suppressed counts run Non-Binary/Undesignated Gender 24,934; American Indian or Alaska Native 23,366; Male 14,507; Female 14,507; Asian/NHPI 14,068; Two or More Races 13,388; Black or African American 9,920; White 5,614; Economically Disadvantaged 5,232; Hispanic/Latino 4,572; Students with Disabilities 4,202; All Students 0. Non-Binary carries 14 reported values of 25,308 (0.055%) and American Indian 194 (0.767%).

The misread: Rendering only the cells that were published. Because the holes are patterned by group size, a chart of the surviving cells redraws discipline as a White, Hispanic, Black, economically-disadvantaged and disabled-student phenomenon — those being the groups whose cells survive. Equally: computing any ratio, risk index or 'X times more likely' statement between two groups censored at different rates.

Full entry, with the story and the numbers, in the served ergo doc.

published-zero-is-not-a-gap [corrupts · mitigated] `0` and `0.00%` are exact publisher zeros, not missing values, and they are the majority state in the early grades and the small districts

Type: suppression — applies to tables: discipline_safety_district_group_grade · columns: value_state, reported_literal · rows: value_state IN ('exact-reconciliation-zero','suppressed')

How to spot it: The district student-group and grade slice holds 325,103 exact-reconciliation-zero cells (`0` 154,130 and `0.00%` 170,973) against 183,192 suppressed and 62,047 reported. Grade PK is 78.6% published zero and 0.33% published rate. Across the whole district level, 97.733% of cells in districts under 250 students are either a suppression marker or a published zero, and the Pearson r between log10 enrollment and district suppression rate is only -0.2756 (n=665) — suppression is near-flat at 29-31% above 250 students while the reported rate rises 13.7x across the range.

The misread: Rendering a published zero as 'no data', or a `*`/`<5` as zero. The two errors reverse a finding in opposite directions. Also: writing that small districts are censored more heavily — they are not; they are recorded as zero rather than published as a number.

Full entry, with the story and the numbers, in the served ergo doc.

dimensions-effectively-never-published [misleads · mitigated] Non-Binary/Undesignated Gender, American Indian or Alaska Native and Grade PK are present in every district row and almost never carry a value

Type: coverage — applies to tables: discipline_safety_district_group_grade · rows: conformed_dimension_label IN ('Non-Binary/Undesignated Gender','American Indian or Alaska Native','Grade PK')

How to spot it: District-level reported cells: Non-Binary/Undesignated Gender 14 of 25,308 (0.055%), and 34 of 120,650 across all organization levels (0.028%); American Indian or Alaska Native 194 of 25,308 (0.767%); Grade PK 22 of 19,228 (0.114%). On the three-measure module surface Non-Binary is withheld in 98.05% of cells and American Indian in 90.24%.

The misread: Reading these rows as zero, or as evidence that these students are never disciplined. The rows are withheld, not empty.

Full entry, with the story and the numbers, in the served ergo doc.

arrest-columns-are-named-police [misleads · mitigated] The source fields behind the arrest measures are named PoliceCount and PolicePercent while the published layout defines them as all incidents leading to arrest

Type: definitional — applies to tables: discipline_safety_metric_catalog, discipline_safety_district_group_grade · columns: publisher_defect · rows: metric_id LIKE 'spr.discipline.arrests.any-%' OR metric_id = 'spr.discipline.police-notifications.unique-count'

How to spot it: metric_definitions carries publisher_defect = 'Source field says Police; layout defines all incidents leading to arrest.' on spr.discipline.arrests.any-count and spr.discipline.arrests.any-percent, and 'Source field is OtherIncidents; layout defines a total unique count.' on spr.discipline.police-notifications.unique-count.

The misread: Reconciling the arrest column against the separate police-notification column, or treating one as a subset of the other. They are different tabs with differently-named source fields, and the site never adds or nets them.

Full entry, with the story and the numbers, in the served ergo doc.

reported-incidents-measure-reporting [misleads · open] Every measure here is an administrative record of what a district recorded and referred, so a diligent recorder looks worse than a district that handles the same behaviour informally

Type: definitional — applies to tables: discipline_safety_district_group_grade, discipline_safety_district_headline, discipline_safety_district_all_students

How to spot it: All 54 metrics carry value_origin = 'reported' and universe text describing publisher-reported counts and percentages. There is no independent incident audit in the file and no adjustment for reporting practice.

The misread: Reading any of these numbers as behaviour, crime, safety, school quality or the effect of a budget decision; combining removals, police notifications and arrests into a single index or safety score; or ranking districts on any of them.

Full entry, with the story and the numbers, in the served ergo doc.

Changelog

  • — Widened the district projection from the three never-suppressed slices to every student group and every grade (49,950 to 620,292 consumer rows), carrying all 183,192 district-level suppressions into the warehouse for the first time; added a 54-row metric catalog; asserted the three pre-existing marts byte-identical; shipped the district-profile module and registered the ten issues a student-group discipline surface has to get right, including that `<5` bounds a headcount at zero rather than one.

The text on this page was generated by Claude Opus 5, working as part of a stack of tools created by Lyra Forge.