Skip to content

PAN Lab example

Cleveland State remote proctoring

A camera in the bedroom and a flag that lands by skin tone

Cleveland State University required a webcam scan of students' rooms before online exams, and proctoring software flagged suspected cheating. A court held the scan unconstitutional.

See more

Cleveland State University, a public university in Ohio, used remote-proctoring tools from Honorlock and Respondus to supervise online exams. Before an exam, the student panned the webcam around the room so the software and a proctor, a person supervising the exam, could check for notes or people. The software watched the student by webcam and flagged behavior it judged suspicious.

The ruling

A student challenged the room scan required before a spring 2021 chemistry exam. The room was in the student's home. Such a room is typically a bedroom, a kitchen, or a room in a shared apartment.

On 22 August 2022, Judge Calabrese of the US District Court for the Northern District of Ohio decided the case, Ogletree v. Cleveland State University. The court held that scanning the inside of a student's home was a search under the Fourth Amendment, and an unreasonable one. The Fourth Amendment is the part of the US Constitution that protects people against unreasonable searches.

It was a first-of-its-kind ruling that a routine, widely used proctoring practice violated a student's constitutional rights in their own home.

What the ruling changes

A university tends to see remote proctoring as an integrity tool. It keeps a grade meaningful when the exam is not in a supervised room.

The ruling says it is also a surveillance decision with a rights cost. The most invasive part, the room scan of a private home, can be found unlawful whatever the integrity goal. The risk of cheating is real. The surveillance used against it is not a free default.

A study of who gets flagged

A peer-reviewed study by Yoder-Himes and colleagues was published in Frontiers in Education in 2022. It examined proctoring software in four large university courses in science, technology, engineering, and mathematics, in fall 2020.

The software failed to detect faces more often for darker-skinned and Black students. It flagged them more often and gave them higher priority for review. The study found no matching difference in actual cheating.

The extra flags did not catch more cheating. They were errors, raised more often for darker-skinned students. The sources do not tie the study to Cleveland State's tools or students.

Why a flag is not neutral

A proctoring flag marks something for a proctor to look at. If staff pursue it, the student has to answer it, through a review, a challenge, or sometimes a charge. So a flag rate that is higher for darker-skinned students, with no more actual cheating, is a burden of suspicion distributed by race. Surveillance like this does not treat every student the same, even when every student is equally innocent.

What can be weighed and measured

Each cost can be weighed or measured before students must open their homes to a camera.

The rights cost is a question of proportionality: does the risk of cheating justify this intrusion, or would a less invasive way give the same assurance? The unequal burden can be measured directly, as the flag rate for each group against actual misconduct.

Neither cost shows in the tool's own measure of success, which counts flags raised. The sources describe neither check at Cleveland State before it required the room scan.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network, where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it.

This case has a budget of 9 units. Explore (No Targets) sets no targets. There, and under Service Targets Only, one tool is enough to keep the mistakes on this network contained: Escalate checks, at 2 units.

Under Service Targets Only, Escalate checks also meets the service target, because the software keeps clearly helping the work. Other routes exist, such as Gate record entries with Review on schedule, at 5 units.

Not every route that contains the mistakes meets the service target. Pause AI on alarms stops proctors receiving flags while it holds. Paired with Escalate checks, it contains the mistakes, but so much of the software's help is lost that the service target is missed.

Under Service and Safety Targets and All Governance Targets, the targets are not fully addressable with the available tools. Within the budget, there are 196 allowed combinations of one or more of the eight tools and their settings. None meets every target under either level.

The budget is not what stands in the way. With the budget set aside and every tool on, each at its stronger setting where it has one, four failure pathways stay open.

They are the room scan taken into the session, the session video the software analyzes, the recordings proctors review, and the proctors' review of each flag. No tool offered here closes any of them, because they are how the proctoring does its work.

No tool offered here adds the proportionality review either. A court answered that question afterward. This is a finding about the deployment, not a gap in your approach.

Stylized model of a documented deploymentEducation AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Proctoring-surveillance-class with two established costs network: 6 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 5 assumptions
  • assumed

    This network draws the room scan as its own input, because it is what the court's ruling was about. The scan is not part of watching the exam. It takes in whatever is in a room of the student's home before the exam starts. The network also draws two reviews of the room scan. The university's own proportionality review could have run before any exam, and the sources describe none. The federal court did rule, after one student brought a lawsuit, and its answer came after the exam. The network assumes the workload is a group of students taking an exam. It assumes the flags are read by instructors who review them alongside their teaching.

  • baseline

    This network models the pattern the case file documents. It is not a copy of the real system. Cleveland State University required students to scan their room by webcam before an online exam. Remote-proctoring software then flagged suspected cheating from the video. A federal court held the room scan an unreasonable search under the Fourth Amendment. That is a decided ruling, and the network states it as one. Separately, a peer-reviewed study of proctoring software in four large university courses found more face-detection failures for darker-skinned and Black students. The software also flagged them more often and gave them higher priority for review. The study found no matching difference in actual cheating.

  • baseline

    The network draws the unequal burden of flags as an audit the sources do not show being run. That audit would compare the flag rate for each group of students against actual misconduct, before flags become accusations. The disparity itself comes from the outside peer-reviewed study. The network records it and never calculates it. A flag that staff pursue is an accusation the student must answer. So a higher flag rate for one group, with equal innocence, is a heavier burden of suspicion on that group. The session video the software analyzes is marked sensitive for privacy, and the study found the face-detection failures in video of this kind.

  • assumed

    The network draws the rights cost as a proportionality review the sources do not show being run. That review would weigh how invasive the surveillance is, a room scan of a home, against the risk of cheating. It would also ask whether a less invasive means would do. A court drew that line afterward, holding the room scan an unreasonable search. So surveillance for exam integrity is not a free default. It is a decision with a rights cost, which a court can rule on whatever the integrity goal. The proportionality question is owed before the surveillance is required. Three pathways are marked sensitive for privacy: the room scan taken into the session, the video saved to the record, and the video the software analyzes.

  • assumed

    Nothing in this network works out an outcome for any student. It traces how mistakes pass between the software, the proctors and instructors, the exam record, and the two reviews. The students being watched are outside its workings. The room scan ruling, the unequal flag rates, and the proportionality question are described in the case file. None of them is calculated from anything in this network.

What this example does not show

Show all 2 limitations
  • No student outcome is worked out here. The network traces how mistakes pass between the university's software, staff, and records. The students being watched are outside its workings. The room scan ruling, the unequal flag rates, and the proportionality question are described in the case file, and nothing in the network calculates them.
  • The ruling is a decided federal decision, so its holding is stated as fact: the room scan was an unreasonable search. The unequal flag rates come from a peer-reviewed study of four large university courses, recorded as an outside finding and not calculated here. The flag-rate audit and the proportionality review are included as checks the sources do not show in use.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • A public university required students to pan their webcam around their home before an online exam, using remote-proctoring software that flags suspected cheating from the video. A federal court held that the pre-exam room scan was an unreasonable search under the Fourth Amendment — a first-of-its-kind ruling that a routine proctoring practice violated a student's constitutional rights in their own home. Separately, peer-reviewed measurement of automated proctoring found proctoring software produced more face-detection failures, more red flags, and higher priority scores for darker-skinned and Black students, with no corresponding difference in actual cheating. The deployment is the education domain's clearest case of surveillance-based integrity AI whose costs — a rights violation and a demographic burden of suspicion — are each independently established.

    empirical
    • Regulatory Ogletree v. Cleveland State University, No. 1:21-cv-00500 (N.D. Ohio, August 22, 2022); via Higher Ed Dive and Future of Privacy Forum analyses. https://fpf.org/blog/federal-court-deems-universitys-use-of-room-scans-within-the-home-unconstitutional/
    • Peer-reviewed Yoder-Himes, D.R., Asif, A., Kinney, K., et al. (2022). Racial, skin tone, and sex disparities in automated proctoring software. Frontiers in Education, 7, 881449. https://doi.org/10.3389/feduc.2022.881449 https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2022.881449/full
  • The lesson the case carries is that surveillance-based integrity AI is not a free default: it carries a rights cost that can be independently adjudicated and a demographic burden that can be measured, and both are owed a reckoning before the surveillance is imposed, not after a court or an audit finds the harm. A room scan of a student's home was held to be an unreasonable search, so the surveillance has a rights dimension a court can rule on regardless of the integrity goal. And because peer-reviewed measurement found proctoring software flagging darker-skinned and Black students more often with no more actual cheating, and a flag is an accusation the student must answer, a disparate flag rate is a disparate burden of suspicion. The governable surfaces are the proportionality of the surveillance to the integrity problem it is trying to solve, and the measured flag rate by group.

    empirical
    • Peer-reviewed Yoder-Himes, D.R., Asif, A., Kinney, K., et al. (2022). Racial, skin tone, and sex disparities in automated proctoring software. Frontiers in Education, 7, 881449. https://doi.org/10.3389/feduc.2022.881449 https://www.frontiersin.org/journals/education/articles/10.3389/feduc.2022.881449/full
    • Regulatory Ogletree v. Cleveland State University, No. 1:21-cv-00500 (N.D. Ohio, August 22, 2022); via Higher Ed Dive and Future of Privacy Forum analyses. https://fpf.org/blog/federal-court-deems-universitys-use-of-room-scans-within-the-home-unconstitutional/

Where this connects

Institutional pressures in this domain

  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.

All of them in context on the Education AI domain page.

Levers available here and the patterns behind them

Documented case histories