Skip to content

PAN Lab example

Wisconsin DEWS

Wrong most of the time and unevenly by race

Wisconsin's Dropout Early Warning System labeled students 'high risk'. An audit found most graduated anyway, and Black and Hispanic students were wrongly labeled more often.

See more

The Dropout Early Warning System, or DEWS, was built and run by the Wisconsin Department of Public Instruction. It combined several machine-learning models into one score of each grade 6 to 9 student's risk of not graduating on time. Students with the highest-risk scores were labeled 'high risk' on dashboards school staff saw from 2012 to 2023.

How it was used

The department's analyst who built DEWS published a detailed account of its design in 2015. Its goal was to find students at risk early enough to help them.

The department's own page on the dashboards described the score as a starting point for inquiry. Administrators in a survey of 80 districts reported no training on how to interpret a 'high risk' label. They had no guidance on what it did or did not mean, or what to do about it. About 38 percent of the districts that answered the survey used DEWS.

The case file describes the label as tied to a dashboard, not to a program of help with staff and funding behind it.

What the audit found

In April 2023, The Markup, with Chalkbeat, published an independent audit covering about a decade of the system. When DEWS predicted a student would not graduate, it was wrong about 74 percent of the time. Most students it labeled 'high risk' graduated anyway.

A false alarm is a 'high risk' label on a student who went on to graduate. False alarms happened at a higher rate for Black and Hispanic students than for white students. The rate was 42 percentage points higher for Black students and 18 points higher for Hispanic students. The Markup also reported that the system used race and income to label students. The department had conducted internal equity research on the system and had not published it.

The department's own equity-analysis function found the system unfair in 2021. Its finding was not acted on for nearly two years.

How it ended

The department stopped offering the DEWS dashboard data on 12 October 2023. It said it was evaluating the future of these early warning systems.

What the case shows

A risk label is not help in itself. It is a signal given to a person, who then does something with it, or nothing.

The case file reads this label as wrong most of the time, uneven by race, and handed to staff without training. So the most likely thing it changed was how a student was seen, not the help the student got. A 'high risk' label can lead staff to expect less of a student, or to watch the student more closely.

So watch what a label tied to a dashboard, not to help, actually does.

The counter-example

Chicago Public Schools used a different kind of early warning, the Freshman On-Track indicator. It checks whether a ninth grader has enough credits and no more than one failing grade in a core course. It came from University of Chicago research.

The district paired the indicator with a real program of help for the students it identified. Its graduation rate rose to a record 85 percent in 2023.

It bears on this case because it is the same kind of early warning, built on a simple measure staff can understand and tied to help. The case file places the benefit in that help, not in a more sophisticated prediction.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. Here a mistake is, for example, a student wrongly labeled 'high risk', or a staff note written from that label. A pathway counts as closed once it passes on only a few mistakes. It need not stop them all.

This case starts with one pressure on, AI-literacy gap widens, because administrators reported no training on the label. It has a budget of 11 units. Understand the system, a tool that lowers other tools' prices, is not among its tools, so every price here is full.

Explore (No Targets) sets no targets. There, three pairs of tools, each costing 4 units, are enough for mistakes to die out instead of spreading. Under Service Targets Only, the same three pairs also meet the targets. Those targets include DEWS helping the work it was built for: finding students early enough to help them. One is Escalate checks with Mark AI-written records.

In that pair, Escalate checks stops mistakes passing along the 'high risk' label to staff. Mark AI-written records stops them passing along two pathways: student data scored by DEWS, and labels staff read on the dashboard.

Within the budget, 166 different sets of tools meet the targets under Service Targets Only, counting each tool at either setting. That count uses the Dynamics setting Both, where it starts, with side effects and lingering effects on. With Dynamics off, 174 sets do.

Service and Safety Targets and All Governance Targets are not fully addressable with the available tools. Both levels ask you to close every failure pathway. No tool here acts on the pathway where staff act on the label, so it stays open. It stays open even with every tool at its strongest setting.

More is not better here. Every tool at its strongest setting at once costs 42 units, far over the budget. It also leaves DEWS hurting the work.

Stylized model of a documented deploymentEducation AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Early-warning-class whose label became a lens network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 5 assumptions
  • assumed

    The dashboard is drawn as its own part, because it was what this system delivered to a school. Whether any help followed a name on it is a separate question. This network assumes a very large workload and little capacity to interpret it. DEWS scored every student in four grade levels across the state. Administrators reported no training on reading the label. Every student in those grades was scored, and the capacity to interpret the label was never built.

  • baseline

    This network follows the early-warning system documented in the case file. It does not rebuild the actual system. Wisconsin's DEWS scored every grade 6 to 9 student and gave school staff a 'high risk' label for about a decade. The Markup's audit found it wrong about 74 percent of the time when it predicted a student would not graduate. Its false-alarm rate, the share of 'high risk' labels on students who graduated, was higher for Black and Hispanic students. Administrators in a survey of districts reported no training on the label. The state stopped publishing the dashboards in 2023. The audit figures are the investigation's own analysis, entered as reported.

  • assumed

    This network places the group difference in false alarms on the pathway where DEWS scores student data. The model learned from student data including race and income. The difference is a finding the audit recorded, never computed on this network. The pathway from DEWS back to itself stands for the point that a risk label is not an intervention. It becomes help only if a trained person acts on it and a program of help exists. Tied to a dashboard instead, it changes how staff see the student.

  • baseline

    This network draws two checks the department could have had. The first is a published check of the label's accuracy for each group of students. The department ran internal equity research and did not publish it, so the disparity became an outside finding. The second is training staff and attaching help to a label. Administrators reported no training, and the label was tied to a dashboard, not to help. Chicago Public Schools paired a simple, transparent on-track indicator with real help, and graduation rose to a record. The sources conclude that an indicator staff can understand, tied to help, can outperform an opaque model that only labels.

  • assumed

    This network models no student outcome. It follows how mistakes pass between the institution's model, its staff, and its records. The students being scored are outside the network. The case file holds the audit's accuracy and disparity findings, the lack of training, and the unpublished equity research. It also holds the counter-example of a simple indicator tied to help. None is computed from anything in this network.

What this example does not show

Show all 2 limitations
  • This example models no student outcome. It follows how mistakes pass between the institution's model, its staff, and its records. The students being scored are outside the network. The case file holds the audit's accuracy and disparity findings, the lack of training, and the unpublished equity research. It also holds the counter-example of a simple indicator tied to help. None is computed on this network.
  • The 74 percent figure and the higher false-alarm rates for Black and Hispanic students are The Markup's own analysis, entered as reported. This network shows the disparity on the pathway where DEWS scores student data, as a recorded finding, not a computed harm. The equity audit and the training and intervention loop are drawn as two checks the department could build.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A state's statewide dropout early-warning system used ensemble machine learning to label every grade 6 to 9 student's risk of not graduating on time and delivered the label to school staff through dashboards for about a decade. An independent, decade-scale audit found the system was wrong roughly 74 percent of the time when it predicted a student would not graduate, produced higher false-alarm rates for Black and Hispanic students, and that the deployer's own internal equity research had gone unpublished — while a survey of districts found administrators reporting no training on how to interpret a 'high risk' label. The state stopped publishing the dashboards in 2023 and said it was evaluating the system's future. The deployment is the education domain's clearest case of a risk label whose error and group disparity entered how students were seen rather than the help they received.

    empirical
    • Investigative Feathers, T. (2023, April 27). False Alarm: How Wisconsin Uses Race and Income to Label Students 'High Risk'. The Markup (with Chalkbeat). https://themarkup.org/machine-learning/2023/04/27/false-alarm-how-wisconsin-uses-race-and-income-to-label-students-high-risk
    • Peer-reviewed Knowles, J.E. (2015). Of needles and haystacks: Building an accurate statewide dropout early warning system in Wisconsin. Journal of Educational Data Mining, 7(3), 18-67. https://doi.org/10.5281/zenodo.3554725 https://jedm.educationaldatamining.org/index.php/JEDM/article/view/JEDM082
    • Government Wisconsin Department of Public Instruction. WISEdash for Districts: Dropout Early Warning System (DEWS) Dashboards (including the October 12, 2023 retirement notice). https://dpi.wi.gov/wisedash/districts/about-data/dews
  • The lesson the case carries is that a risk label is only as good as the intervention it triggers and the training of the human who reads it. A label that is wrong most of the time, delivered to staff with no guidance on interpreting it, imports the model's error and its group disparity into how students are perceived rather than into a resourced response — the flag becomes a lens on the student rather than a trigger for help. Set against this, a large district's transparent, low-tech on-track indicator, built on interpretable research and paired with real intervention, accompanied a rise in graduation to a record level. The contrast locates the benefit in the intervention the indicator makes legible enough for staff to act on well, not in the sophistication of the prediction — an interpretable indicator that drives help can outperform an opaque model that only labels.

    empirical
    • Academic Allensworth, E.M., & Easton, J.Q. (2007). What Matters for Staying On-Track and Graduating in Chicago Public Schools. University of Chicago Consortium on School Research. https://consortium.uchicago.edu/publications/what-matters-staying-track-and-graduating-chicago-public-schools-focus-students
    • Investigative Feathers, T. (2023, April 27). False Alarm: How Wisconsin Uses Race and Income to Label Students 'High Risk'. The Markup (with Chalkbeat). https://themarkup.org/machine-learning/2023/04/27/false-alarm-how-wisconsin-uses-race-and-income-to-label-students-high-risk

Where this connects

Institutional pressures in this domain

  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.

All of them in context on the Education AI domain page.

Levers available here and the patterns behind them

Documented case histories