PAN Lab example
Unilever and HireVue graduate hiring
Good audits with real savings — and a cohort no one can see
A graduate-hiring pipeline chains a games screen with video-interview scoring. Modeled on a deployment with its audits on the record: ~90% faster hiring, ~£1M saved, +16% diversity - all company- or vendor-reported - and both vendors audited, honestly but partially. This is the audit lever in its good form. So watch the thing even good audits do not reach: every one of those numbers is measured on the people the system advanced, and says nothing about the people it screened out.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Graduate-hiring-class with honest, partial audits network: 5 components and 10 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 3 assumed · 2 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
The record describes a two-stage pipeline built from two different vendors' products - a games-based screen chained to automated video-interview scoring - so the video stage is drawn as its own model fed by the games stage rather than folded into one component. That chain is the structural fact the domain's other deployments lack: the second stage only ever assesses the population the first advanced, so a first-stage skew is invisible downstream no matter how well the second stage is audited. The outcome pathway is drawn at full strength because hire outcomes are the assessment's training signal and the selected cohort is the only cohort that produces them.
- baseline
This models the two-sided graduate-hiring pattern documented in the case file - not a reconstruction of the actual pipeline. The deployer's reported results (~90% faster time-to-hire, ~50,000 interview hours saved, ~£1M annual savings, 16% diversity improvement) are company- or vendor-reported and not independently audited - drawn as the deployer's dashboard, the family's service regime seen from inside, not an external measurement.
- baseline
The audit lever is drawn present but partial, at a low level on the independent model check: both vendors' audit machinery is public in its honest but partial form - a cooperative academic audit of the games vendor with source-code access (its four-fifths-rule de-biasing found faithfully implemented, with the recorded caveat that vendor staff were co-authors), and the video vendor retiring its facial-analysis input under scrutiny (visual features added only ~0.25% predictive power) with a narrow-scope external audit. Real, partial, and scoped - neither absent (the abandonment case) nor shielded (the litigation case).
- assumed
The domain's deepest blind spot is drawn as the latent oversight check, empty at baseline: rejected candidates never re-enter the outcome data, so the reported quality and diversity effects are measured on hires only. No audit quality reaches this, because it is upstream of the audit - a 16% diversity gain among hires is consistent with a diversity loss among those rejected, and the measurement cannot see the rejected cohort at all. The population you can measure is the population you selected.
- assumed
No candidate outcome is modeled here. This Lab reads institutional propagation only, and applicants are boundary-only. The reported dashboard figures, the audits, and the rejected-candidate blind spot live in the case file, and are never computed from anything in this diagram.
What this example does not show
- No candidate outcome is modeled. The Lab reads institutional propagation only; applicants are boundary-only, and the reported dashboard figures, the vendor audits, and the rejected-candidate blind spot live in the case file, never computed on this diagram.
- The ~90%/£1M/+16% figures are company- or vendor-reported dashboard numbers entered as such, not independently audited; the two vendor audits are real but partial (one vendor-co-authored, one narrow-scope), and the rejected-cohort blind spot is drawn as a latent check, not a computed harm.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
A graduate-hiring pipeline chained a games-based assessment with automated video-interview scoring, and the deployer reports roughly a 90 percent reduction in time-to-hire (from about four months to about four weeks), around 50,000 candidate interview hours saved, about one million pounds in annual savings, and a 16 percent improvement in diversity. Every one of those figures is company- or vendor-reported and none is independently audited, so they are the deployer's own dashboard rather than an external measurement — which is exactly what the family's service regime looks like from inside.
empirical- Vendor Best Practice AI. Unilever saved over 50,000 hours in candidate interview time and delivered over £1M annual savings and improved candidate diversity with machine analysis of video-based interviewing (AI case study). https://www.bestpractice.ai/ai-case-study-best-practice/unilever_saved_over_50,000_hours_in_candidate_interview_time_and_delivered_over_%C2%A31m_annual_savings_and_improved_candidate_diversity_with_machine_analysis_of_video-based_interviewing.
Both vendors' audit machinery is on the public record in an honest but partial form. The games vendor underwent a cooperative academic audit with source-code access, in which its four-fifths-rule de-biasing pipeline was found faithfully implemented — with the independence caveat that vendor staff were co-authors — and the video vendor retired its facial-analysis input under scrutiny after internal research found visual features added only about 0.25 percent predictive power, publicizing a narrow-scope external audit. The family's structural blind spot applies in full: rejected candidates never re-enter the outcome data, so the claimed quality and diversity effects are measured on hires only.
empirical- Peer-reviewed Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
- Trade press Maurer, R. (2021). HireVue Discontinues Facial Analysis Screening. SHRM; with HireVue and ORCAA audit announcements (2021). https://orcaarisk.com/in-the-news/2021/1/12/orcaas-audit-of-hirevue-is-live
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
All of them in context on the Hiring & employment screening AI domain page.
Levers available here and the patterns behind them
- Gate vendor updates — Vendor quality gate
- Review on schedule — Oversight cadence & retrospectives
- Check with a second model — Cross-model verification
- Check copied records — Reconcile copied records
- Escalate checks — State-feedback vigilance
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- Graduate-hiring AI with its audits on the record
- A resume screener that learned the past's bias
- Vendor screening across thousands of employers (litigation live)
- HireVue video assessment (vendor layer)
- The 1959 statute and the integrity video screen (Baker v. CVS Health)
- An internal promotion, a recorded screen, and a captioning request (D.K. charges against Intuit and HireVue)
- Aon pre-hire assessment suite (vendor's own tables)
- The cooperative audit: a paid source-code examination, and what happened to its verdict
- McHire and the 64-million-record custody exposure
- SiriusXM's iCIMS applicant screening
- Checkr gig-economy background screening
- The rule with no number to disclose
- The account goes dark at nine; the reason arrives on day twenty-six
- iTutorGroup Tutor Application Screen
- Meta Job-Ad Delivery: the guardrail and the layer below