PAN Lab example
Amazon recruiting engine
It learned who was hired rather than who succeeds
A resume scorer is trained on ten years of the company's own hiring. Modeled on an experimental tool that learned the history rather than the merit: it penalized the word 'women's' and downgraded women's-college graduates. The team found the bias, patched the terms, and hit a wall - the model had learned the pattern, not the words. Watch the two things this case names: training data as a selection function, and the ceiling on patching a model that learned a proxy.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Resume-screener-class trained on the past's hiring network: 6 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 4 assumed · 1 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
Two documented components were drawn that the earlier template omitted. The engineering team that built the scorers in-house is its own group of staff with a retrain-and-patch pathway into the model: this deployment owned the tool outright, unlike the vendor platform and the two-vendor pipeline elsewhere in this domain, which is why improve-model is genuinely this organization's lever. The term-level correction is drawn as a mediating guardrail rather than a phrase in a pathway's copy, because that component is the case - the finding is its ceiling. The training pathway is drawn at full strength as the domain's strongest contamination: ten years of the organization's own hiring are the label set, not merely an input. The discontinuation authority is drawn faint rather than empty because it was exercised - this is the one deployment in the domain where the authority to stop was used.
- baseline
This models the resume-screener-trained-on-past-hires pattern documented in the case file - not a reconstruction of the actual tool. Its defining mechanism is the lock-in (drawn on the record-to-model training edge): a screener trained on ten years of the organization's own hiring decisions imports the past's selection function, learning who was hired rather than who succeeds, and reproducing the history's bias (predominantly male hires; penalized gender proxies) as a prediction.
- assumed
The patch lever has a documented ceiling, drawn as the latent independent model check: removing the named proxies the model used does not remove the pattern it learned, because the model re-derives the signal from unnamed correlates. The lever that actually breaks the loop is a design that values exploring candidates the history under-selected, not a term-patch - a harder and different thing than scrubbing words.
- assumed
Abandonment was the honest exit, drawn as the latent oversight check: the team detected the bias, could not guarantee neutrality against unknown proxies, and scrapped the tool - before any external harm was documented. 'We could not make it fair, so we stopped' is a legitimate governance outcome, and it is the honest end of a ladder this domain otherwise shows shielded or litigated. The record is investigative, not a company publication - part of the honest structure.
- assumed
No candidate outcome is modeled here. This Lab reads institutional propagation only, and applicants are boundary-only. The gender-proxy finding, the patch ceiling, and the abandonment live in the case file, and are never computed from anything in this diagram; the tool was advisory (never a sole ranking) and was scrapped.
What this example does not show
- No candidate outcome is modeled. The Lab reads institutional propagation only; applicants are boundary-only, and the gender-proxy finding, the patch ceiling, and the abandonment live in the case file, never computed on this diagram.
- The record is investigative (reported through interviews with team members), not a company publication; the tool was experimental, advisory, and scrapped, and the diagram draws the lock-in and the patch ceiling as structure, it does not compute bias.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
An internal team built an experimental recruiting engine — roughly 500 models scoring resumes one to five stars per role and location — trained on ten years of the company's own hiring decisions, a period whose hires were predominantly male. The models learned that history: they penalized the word 'women's' and downgraded graduates of women's colleges, reading gender proxies as negative signal. The team patched the identified terms but concluded that term-level fixes could not guarantee neutrality against unknown proxies, because the model had learned the pattern rather than the words, and the company scrapped the project around 2017; per the company, recruiters saw the tool's recommendations but it was never used as a sole ranking.
empirical- Investigative Dastin, J. (2018, October 10). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters. https://www.euronews.com/business/2018/10/10/amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women
Training a screener on an organization's past hiring decisions imports the past's selection function: research on hiring as exploration finds that models trained on prior hires raise hire rates but replicate historical selection, and that a screener which values exploration rather than only exploitation breaks that lock-in loop. Two governance lessons follow — the patch lever has a documented ceiling, since removing named proxies does not remove a learned correlation, and abandonment can itself be a governance outcome, taken here before any external harm was documented rather than after an adjudication.
empirical- Peer-reviewed Li, D., Raymond, L.R., & Bergman, P. (2020). Hiring as Exploration. NBER Working Paper 27736. https://doi.org/10.3386/w27736 https://www.nber.org/papers/w27736
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Hiring & employment screening AI domain page.
Levers available here and the patterns behind them
- Mark AI-written records — Provenance labeling
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Assign a challenger — Structured dissent
- Pause AI on alarms — Deployment circuit-breaker
- Escalate checks — State-feedback vigilance
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- A resume screener that learned the past's bias
- Vendor screening across thousands of employers (litigation live)
- Graduate-hiring AI with its audits on the record
- HireVue video assessment (vendor layer)
- The 1959 statute and the integrity video screen (Baker v. CVS Health)
- An internal promotion, a recorded screen, and a captioning request (D.K. charges against Intuit and HireVue)
- Aon pre-hire assessment suite (vendor's own tables)
- The cooperative audit: a paid source-code examination, and what happened to its verdict
- McHire and the 64-million-record custody exposure
- SiriusXM's iCIMS applicant screening
- Checkr gig-economy background screening
- The rule with no number to disclose
- The account goes dark at nine; the reason arrives on day twenty-six
- iTutorGroup Tutor Application Screen
- Meta Job-Ad Delivery: the guardrail and the layer below