PAN Lab example
Unilever and HireVue graduate hiring
Good audits with real savings — and a cohort no one can see
Unilever screened recent graduates with two vendors' AI tools: pymetrics games, then HireVue video-interview scoring. Its reported gains cover only the applicants it advanced.
See more
Unilever used two vendors' AI tools in its graduate hiring from about 2016. First, the vendor pymetrics screened applicants with a games-based assessment. Then HireVue's software automatically scored the recorded video interviews of those the games advanced, before Unilever's recruiters reviewed them.
How the pipeline worked
Unilever used the pipeline for early-careers hiring worldwide, in a program led from the UK and the US. Unilever chained the two vendors' tools. It set the cutoff between the two stages and the threshold for the shortlist. The video stage's rankings shaped the shortlist.
The games assessment rejected applicants below its cutoff automatically. Recruiters then reviewed the shortlist and ran the final round.
What Unilever reports
Unilever reports that time to hire fell by roughly 90 percent, from about four months to about four weeks. It reports around 50,000 hours of candidate interview time saved, and about £1 million saved a year. It also reports a 16 percent improvement in the diversity of hires. The case study is undated. The pipeline ran from about 2016.
Every one of these figures is reported by the company or a vendor, and none is independently audited. They come from an industry case study by Best Practice AI. They are Unilever's own dashboard: the tools' benefits seen from inside the company, not an outside measurement.
The audits
Both vendors' audits are on the public record, in an honest but partial form.
Academic researchers audited pymetrics with access to its source code, in a study published in 2021. They checked its de-biasing process, which applies the four-fifths rule, a common test that compares how often different groups pass. They found the process faithfully implemented. The study's co-authors included pymetrics employees and the company's chief executive, so its independence is disputed.
HireVue stopped using facial analysis in new assessment models in about March 2020. It announced the change in January 2021, under scrutiny. Its own research had found that visual features made up only about 0.25 percent of its predictive power in most job models.
HireVue commissioned an audit from ORCAA, an auditing firm. It covered one early-career assessment and worked mainly through interviews with stakeholders. It did not examine the tool's design or its training data.
So both audits are real, partial, and scoped. One is a code check co-authored by the vendor. The other dropped an input that barely helped, and had one narrow assessment audited. That is what good auditing looks like in practice, and it is better than most.
Two other hiring cases in this Lab show the forms on either side. Amazon built a résumé-screening engine, found it biased, failed to fix it, and abandoned it. Workday's internal bias testing was shielded from discovery, the exchange of evidence in a lawsuit.
What the reported gains do not show
Unilever records no later outcomes for rejected applicants. The pipeline records what happens to the people it advances and hires, not to the people it screens out. So the claimed quality and diversity effects are measured on hires only.
A 16 percent diversity improvement among those hired says nothing about who was filtered out earlier. It cannot show whether some groups of applicants were screened out more often than others. Diversity among hires could rise while diversity among those rejected falls. A faster, cheaper pipeline that improves the measured group can still be doing unknown things to the group it rejects.
The benefits are real to Unilever and reported in good faith. The audits are real and better than most. Both are measured on a population that leaves out everyone the pipeline turned away.
In a screening system, the people you can measure are the people you selected. The harm a screener does falls on exactly the people who leave no outcome data. So the honest question is not only whether the assessments are fair to the people they advanced. It is what happened to everyone they did not advance, and whether Unilever can even see them.
What this network is drawn from
This network follows the pattern the case file describes. It does not reconstruct Unilever's actual pipeline. It shows the games stage, the video stage, recruiters, the hiring record, and the audits and results measurement. It also shows a check the sources do not describe: a comparison of outcomes for rejected applicants and hires. Applicants themselves are outside the network.
What the available tools can and cannot address
A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. A tool closes a pathway when mistakes stop passing along it. The work along it may go on.
This case has a budget of 10 units. It starts with one pressure, Silent vendor update: a vendor changes its assessment without notice. You can remove it only under Explore (No Targets), which sets no targets.
The other three levels all ask for the network's mistakes to be contained, meaning corrected rather than building on each other. They also ask the assessments to bring enough benefit to the work.
Under Service Targets Only, the targets can be met. The cheapest way costs 5 units: Gate vendor updates at its stronger setting, with Escalate checks. Within the budget, 15 tool sets meet them, or 35 ways once stronger settings are counted. Every set uses Gate vendor updates or Escalate checks, and 8 of the 15 use both.
The Dynamics menu can turn off lingering effects (damage that outlasts its cause). With them off, Check with a second model with Escalate checks also meets the targets for 5 units. Then 19 tool sets meet them, or 20 with side effects off as well.
Service and Safety Targets also asks you to close every failure pathway and keep up with the work, among other targets. Six pathways are open before any tool is used. The first three are Advanced applicants sent to recruiters, Recruiters judge advanced applicants, and Hiring decisions recorded. The other three are Applicant data used by the assessment, Applicant history read, and Games stage gates the video stage.
The offered tools can close three of them. Escalate checks closes Advanced applicants sent to recruiters. Store less data closes Hiring decisions recorded. Check with a second model closes Games stage gates the video stage.
No offered tool closes the other three: Recruiters judge advanced applicants, Applicant data used by the assessment, and Applicant history read. All Governance Targets asks for everything Service and Safety Targets asks, and more. So neither of those two levels is fully addressable with the available tools, even with the budget lifted.
No offered tool adds Outcomes of rejected applicants, the check this case turns on, either.
More is not better here. Using every tool at its strongest setting costs 30 units, well over the budget of 10. It contains the mistakes. But three pathways stay open, and the benefit the assessments bring falls below what the targets ask. So it meets the targets at none of the three levels.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Graduate-hiring-class with honest, partial audits network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
Show all 5 assumptions
- assumed
The sources describe a two-stage pipeline built from two vendors' products: pymetrics games, then HireVue's automated video-interview scoring. So the network draws the video stage as its own part, which receives applicants from the games stage. The second stage assesses only the people the first advanced. A bias in the first stage against some group of applicants stays invisible downstream, however well the second stage is audited. The network assumes hire outcomes are what the assessment learns from. Only the people it selected produce those outcomes.
- baseline
This network follows the graduate-hiring pattern documented in the case file. It does not reconstruct the actual pipeline. Unilever reports a cut of roughly 90 percent in time to hire and about 50,000 hours of candidate interview time saved. It also reports about £1 million saved a year and a 16 percent diversity improvement. These figures are reported by the company or a vendor and are not independently audited. The network treats them as Unilever's own dashboard, its benefits seen from inside, not an outside measurement.
- baseline
The network draws the vendors' audits as a check on the assessments that is present but partial. Both audits are public in an honest but partial form. Academic researchers audited pymetrics with access to its source code. They found its de-biasing faithfully implemented, with the recorded caveat that pymetrics staff were co-authors. That de-biasing applies the four-fifths rule, a common test that compares how often different groups pass. HireVue stopped using its facial-analysis input under scrutiny, after visual features were found to make up only about 0.25 percent of its predictive power in most job models. It commissioned a narrow audit. The audits are real, partial, and scoped. Amazon abandoned a résumé-screening engine it could not fix, and Workday's bias testing was shielded from view in a lawsuit. These audits were neither abandoned nor hidden.
- assumed
The network draws the case's deepest gap as a check on rejected applicants that the sources describe no one running. Unilever records no later outcomes for rejected applicants, so the reported quality and diversity effects are measured on hires only. No audit, however good, reaches this, because the outcome data any audit could check holds no rejected applicants. A 16 percent diversity gain among hires cannot show whether some groups of applicants were screened out more often than others. Diversity among hires could rise while diversity among those rejected falls. The measurement cannot see the rejected group at all. The people you can measure are the people you selected.
- assumed
No applicant's outcome is modeled here. This network shows how mistakes pass between the assessments, recruiters, and the hiring record, and applicants stay outside it. The reported dashboard figures, the audits, and the gap for rejected applicants are recorded in the case file. They are never computed from anything in this network.
What this example does not show
Show all 2 limitations
- This example does not show any applicant's outcome. It shows how mistakes pass between the assessments, recruiters, and the hiring record, and applicants stay outside it. The reported dashboard figures, the vendor audits, and the gap for rejected applicants are recorded in the case file, never computed here.
- This example does not verify Unilever's reported results. The cut of roughly 90 percent in time to hire, about £1 million saved a year, and 16 percent diversity gain are company or vendor figures. They are entered as such and are not independently audited. The two vendor audits are real but partial: pymetrics staff co-authored one, and the other was narrow in scope. The gap for rejected applicants is drawn as a check the sources describe no one running, not as a computed harm.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A graduate-hiring pipeline chained a games-based assessment with automated video-interview scoring, and the deployer reports roughly a 90 percent reduction in time-to-hire (from about four months to about four weeks), around 50,000 candidate interview hours saved, about one million pounds in annual savings, and a 16 percent improvement in diversity. Every one of those figures is company- or vendor-reported and none is independently audited, so they are the deployer's own dashboard rather than an external measurement — which is exactly what the family's service regime looks like from inside.
empirical- Vendor Best Practice AI. Unilever saved over 50,000 hours in candidate interview time and delivered over £1M annual savings and improved candidate diversity with machine analysis of video-based interviewing (AI case study). https://www.bestpractice.ai/ai-case-study-best-practice/unilever_saved_over_50,000_hours_in_candidate_interview_time_and_delivered_over_%C2%A31m_annual_savings_and_improved_candidate_diversity_with_machine_analysis_of_video-based_interviewing.
Both vendors' audit machinery is on the public record in an honest but partial form. The games vendor underwent a cooperative academic audit with source-code access, in which its four-fifths-rule de-biasing pipeline was found faithfully implemented — with the independence caveat that vendor staff were co-authors — and the video vendor retired its facial-analysis input under scrutiny after internal research found visual features added only about 0.25 percent predictive power, publicizing a narrow-scope external audit. The family's structural blind spot applies in full: rejected candidates never re-enter the outcome data, so the claimed quality and diversity effects are measured on hires only.
empirical- Peer-reviewed Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
- Trade press Maurer, R. (2021). HireVue Discontinues Facial Analysis Screening. SHRM; with HireVue and ORCAA audit announcements (2021). https://orcaarisk.com/in-the-news/2021/1/12/orcaas-audit-of-hirevue-is-live
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Hiring & employment screening AI domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- Graduate-hiring AI with its audits on the record
- A resume screener that learned the past's bias
- Vendor screening across thousands of employers (litigation live)
- HireVue video assessment (vendor layer)
- The 1959 statute and the integrity video screen (Baker v. CVS Health)
- An internal promotion, a recorded screen, and a captioning request (D.K. charges against Intuit and HireVue)
- Aon pre-hire assessment suite (vendor's own tables)
- The cooperative audit: a paid source-code examination, and what happened to its verdict
- McHire and the 64-million-record custody exposure
- SiriusXM's iCIMS applicant screening
- Checkr gig-economy background screening
- The rule with no number to disclose
- The account goes dark at nine; the reason arrives on day twenty-six
- iTutorGroup Tutor Application Screen
- Meta Job-Ad Delivery: the guardrail and the layer below