Skip to content

PAN Lab example

Amazon recruiting engine

It learned who was hired rather than who succeeds

Amazon's experimental resume scorer, trained on ten years of resumes mostly from men, penalized "women's". Amazon scrapped it when edits could not guarantee a fix.

See more

Amazon's recruiting engine was experimental software, built by an Amazon engineering team in Edinburgh from about 2014. It was about 500 models that scored resumes from one to five stars, by job role and location. The models learned from patterns in resumes sent to Amazon over ten years.

What went wrong

Most of the resumes Amazon received in those ten years came from men. The models learned that history rather than merit. They penalized the word "women's" and downgraded graduates of women's colleges. A gender proxy is a detail that stands in for gender without naming it. The models treated such proxies as marks against an applicant.

The case file calls this the field's defining mechanism in its clearest form. A screener trained on past hiring decisions does not learn who will succeed. It learns who was hired. It then repeats the rule behind those past choices as if it were a prediction.

What the team did

The team found the bias and edited the models to treat the flagged terms neutrally. Then it reached what the case file calls the correct and important conclusion. Edits to named terms could not guarantee neutrality against proxies no one had found. The models had learned the pattern, not the words.

Remove "women's", and a model can still find the same signal in a dozen details no one has named. The case file calls this the ceiling on patching. An edit covers the proxies you can see, and the learned pattern works around them.

What research on hiring shows

The economists Danielle Li, Lindsey Raymond, and Peter Bergman studied resume screening at a Fortune 500 professional-services firm, in a 2020 paper. It shows the same trap outside Amazon. Screeners trained on past hires raised hire rates but repeated past selection. They picked far fewer Black and Hispanic applicants.

A screener that gave extra weight to applicants past hiring had passed over, to learn how they did improved both hiring quality and diversity. The case file says that kind of design is what breaks that pattern.

How it ended

Amazon scrapped the project and disbanded the team around 2017. According to people familiar with the project, recruiters saw the engine's recommendations but never relied on them alone. Amazon declined to comment.

The case file reads stopping as a governance outcome in its own right. Amazon stopped because a fix could not be guaranteed. It stopped before any outside harm was documented. That sets this case apart from hiring cases that ended in legal proceedings. The case file calls it the cleanest case of an organization finding its own model's bias and choosing to stop.

Where the account comes from

Jeffrey Dastin of Reuters reported the story on October 10, 2018, from five unnamed people familiar with the project. Amazon declined to comment and did not publish its own account. So the clearest account of a company finding its own model's bias comes from reporting, not disclosure.

What this case asks

The case file draws three lessons. First, training data works as a rule for choosing people. "Trained on our own hiring data" is a decision to turn the past's biases into the future's filter. The more skewed that past hiring was, the more closely the model repeats the skew.

Second, editing out terms has a ceiling. Third, "we could not make it fair, so we stopped" is sometimes the correct answer.

What the available tools can and cannot address

A failure pathway is a link between two parts of the network where a mistake made by one part can be passed on to the other. When a tool closes a failure pathway, mistakes stop passing along it, though the link stays in use.

Explore (No Targets) sets no targets. Under Service Targets Only, the targets can be met. The cheapest ways cost 5 of this case's 8 budget units. Each uses Escalate checks, with either Check with a second model or Mark AI-written records at its stronger setting. In all, 26 different sets of tools fit the budget and meet these targets, counting stronger settings.

Under Service and Safety Targets and under All Governance Targets, this case is not fully addressable with the available tools. Those levels ask you to close every failure pathway, among other targets. Three stay open whatever you choose: Past resumes used to train the engine, Recruiters' judgment on scored resumes, and Team edits to the models. All 174 combinations of tools and settings that fit the budget were checked, and each leaves those three open. Lifting the budget does not change that.

No tool here changes what the engine learned from. The case file names a different design as the fix, one that gives extra weight to applicants past hiring passed over, to learn how they do. No tool offered here builds it.

Stylized model of a documented deploymentHiring & employment screening AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Resume-screener-class trained on the past's hiring network: 6 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

Show all 5 assumptions
  • assumed

    Two parts the sources document are included here. One is Amazon's engineering team in Edinburgh, which built the models and edited them. Amazon built the engine itself, with no outside vendor. So upgrading the model was a choice Amazon itself held. Some other hiring cases in the Lab run on vendors' tools instead. The other part is the edit to the flagged terms, included because its limit is this case's finding. On the pathway from the hiring record, ten years of resumes sent to Amazon were the examples the engine learned from. They were not one input among many. The authority to stop is included because it was used.

  • baseline

    This example follows the pattern the case file documents: a resume screener trained on a company's own past hiring data. It is not a copy of Amazon's actual engine. Its central feature is the pathway from the hiring record to the engine. The engine learned from ten years of resumes sent to Amazon, most of them from men. The case file reads this as learning who was hired, rather than who went on to succeed. The engine repeated that skew as a prediction and penalized gender proxies, details that stand in for gender.

  • assumed

    Editing the flagged terms has a documented limit, shown here as a test of the engine for other proxies. Removing the terms the models used does not remove the pattern they learned. The models can find the same signal in details no one has named. The case file says the fix that breaks the pattern is a different design. That design gives extra weight to applicants past hiring passed over, to learn how they do. It is harder than removing words.

  • assumed

    Stopping was the honest way out, shown here as the bias finding and authority to stop. The team found the bias and could not guarantee neutrality against proxies no one had found. The company scrapped the tool before any outside harm was documented. The case file calls "we could not make it fair, so we stopped" a legitimate governance outcome. The case file contrasts this with hiring cases that end with testing kept from outside view, or in court. What is known here comes from reporting, not from anything Amazon published.

  • assumed

    This example does not model what happened to any applicant. It shows how mistakes move among the engine, Amazon's staff, and the hiring record. Applicants are outside the network. The gender-proxy finding, the limit of the edits, and the decision to stop come from the case file. Nothing in this network computes them. According to people familiar with the project, recruiters never relied on the tool alone. It was scrapped.

What this example does not show

Show all 2 limitations
  • This example does not show what happened to any applicant. It shows how mistakes move among the engine, Amazon's staff, and the hiring record. Applicants are outside the network. The gender-proxy finding, the limit of the edits, and the decision to stop come from the case file. Nothing in the network computes them.
  • What is known comes from Reuters reporting based on five unnamed people familiar with the project, not from anything Amazon published. The tool was experimental and was scrapped. According to people familiar with the project, recruiters never relied on it alone. This example shows how learning from past resumes locks in their pattern, and where the edits stop working. It does not measure the bias.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • An internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role and location — trained on patterns in resumes submitted to the company over ten years, most of them from men. The models learned that history: they penalized the word 'women's' and downgraded graduates of women's colleges, reading gender proxies as negative signal. The team patched the identified terms but concluded that term-level fixes could not guarantee neutrality against unknown proxies, because the model had learned the pattern rather than the words, and the company scrapped the project around 2017; according to people familiar with the effort, recruiters saw the tool's recommendations but never relied solely on its rankings.

    empirical
    • Investigative Dastin, J. (2018, October 10). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters. https://www.euronews.com/business/2018/10/10/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women
  • Training a screener on an organization's past hiring decisions imports the past's selection function: research on hiring as exploration finds that models trained on prior hires raise hire rates but replicate historical selection, and that a screener which values exploration rather than only exploitation breaks that lock-in loop. Two governance lessons follow — the patch lever has a documented ceiling, since removing named proxies does not remove a learned correlation, and abandonment can itself be a governance outcome, taken here before any external harm was documented rather than after an adjudication.

    empirical
    • Peer-reviewed Li, D., Raymond, L.R., & Bergman, P. (2020). Hiring as Exploration. NBER Working Paper 27736. https://doi.org/10.3386/w27736 https://www.nber.org/papers/w27736

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Hiring & employment screening AI domain page.

Levers available here and the patterns behind them

Documented case histories