Domain Atlas / Hiring & employment screening AI
HireVue video assessment (vendor layer)
Explore this deployment in the PAN Lab ↗
In the PAN Lab, the readouts of this case's model organization carry a shaded evidence band whose width follows the least-established class among the modeling inputs the readings rest on.
The least-established input behind this case's model organization's readings is an assumption, not a measurement. Evidence base: 1 assumed · 8 published baseline.
HireVue's video-assessment platform is a vendor-layer deployment: one scoring engine behind hundreds of separate employers' hiring pipelines. Candidates record answers to a structured question set, or play game-based assessments; per-assessment models score verbal and paraverbal features of the responses into competency scores that place each applicant in a Bottom, Middle or Top tier, and the client employer chooses where to cut — one deploying employer's published bias audit records it evaluating at Top-plus-Middle against Bottom. The client also chooses whether to use algorithmic scoring at all: as of January 2021 the vendor reported that approximately 20 percent of its customers used the predictive-analytics feature and that the rest used the platform for human review of recorded video. Every scale figure is vendor-reported and none is independently audited: more than 19 million video interviews and more than 700 customers as of January 2021, and more than 33 million interviews, 200 million chat-based candidate engagements and more than 800 customers as of January 2023. The models are built on historical applicant data pooled across employer implementations, and the mandated bias audits are computed on that same pooled store. A rejected candidate generates no outcome data anywhere in that loop, so the store that trains and audits the models cannot observe the population the models screened out.[4]
What happened
A candidate applies for a job and is asked to record video answers to a fixed set of questions, or to play a short set of assessment games. The recording goes to HireVue, a vendor whose platform sits behind hundreds of separate employers' hiring pipelines. Where the employer has switched the scoring feature on, per-assessment models score verbal and paraverbal features of the answers into competency scores and place the candidate in a Bottom, Middle or Top tier. The vendor emits the tier. The employer chooses where to cut: one deploying employer's published bias audit records it evaluating at Top-plus-Middle against Bottom.
Two numbers set the scale, and both are the vendor's own. As of January 2021 it reported more than 19 million video interviews hosted and more than 700 customers; as of January 2023, more than 33 million interviews, 200 million chat-based candidate engagements and more than 800 customers. A third number sets the shape: approximately 20 percent of its customers used the predictive-analytics scoring feature as of January 2021, and the rest used the platform for human review of recorded video. So the engine reaches a very large population, a minority-but-large slice of it is actually scored, and the decision about which slice belongs to the buyer, not to the builder.
On 6 November 2019 the Electronic Privacy Information Center filed a complaint with the Federal Trade Commission. It alleged that assessments marketed as measuring cognitive ability, psychological traits, emotional intelligence and social aptitudes were unfair and deceptive practices under Section 5 of the FTC Act, that the company had falsely denied using facial recognition, and that the results were biased, unprovable and not replicable. No public FTC enforcement action against the company is on the record as of August 2026. The honest description of this channel is that the vendor was the subject of an advocacy complaint; "under regulator scrutiny" overstates what happened.
What happened next is the part of this record worth reading twice, and the order matters. In early 2020 — the Society for Human Resource Management reports the discontinuation as March 2020 — the company stopped including visual analysis in new assessment models. It said nothing publicly for about ten months. On 12 January 2021 it announced the removal, together with the results of an algorithmic audit it had commissioned from O'Neil Risk Consulting and Algorithmic Auditing. Its stated reason came with arithmetic attached: internal research put the contribution of nonverbal visual data at about 0.25 percent of the model's predictive power in most job models, and about 4 percent for high-customer-contact roles. Chief executive Kevin Parker put the rest plainly — it was not worth the concern it was causing people. Both figures are vendor-reported and neither has been independently audited. Whether the 2019 complaint caused the removal is an inference nobody in this record makes; what the record supports is that the measured contribution had fallen to near nothing and public concern had risen, and that a single party could and did close the channel across every deployment at once. No individual buyer held that control.
The audit announced beside the removal is the other half of the January 2021 story, and its perimeter is the honesty core of the whole file. MIT Technology Review's February 2021 examination recorded what it covered and what it did not: one representative pre-built assessment use case in early-career and campus hiring; no examination of the tool's technical design or its training data; structured stakeholder interviews as the method; and a report published on HireVue's own site only behind a nondisclosure agreement. The company's summary — that its assessments work as advertised with regard to fairness and bias — is its own characterization of exactly that scope. Brookings' Alex Engler publicly contrasted its depth unfavourably with a contemporaneous audit of a competitor that reached the source code; legal scholar Pauline Kim noted the conflict risk in an auditor paid by the audited party. Neither audit addressed whether the products improve hiring at all. The audit did produce recommendations, and they were specific: investigate accent bias, and look at the flagging of minority candidates who give brief answers.
Then the law arrived, and it arrived at the buyer. Illinois' Artificial Intelligence Video Interview Act took effect on 1 January 2020. It requires an employer using AI analysis of video interviews to notify the applicant before the interview, explain how the AI works and what general types of characteristics it uses to evaluate them, and obtain consent to be evaluated; it restricts sharing the video, requires deletion within 30 days of a request, and from 2022 requires employers relying solely on AI analysis to report applicant race and ethnicity data annually. Its text states no express enforcement mechanism and no private right of action. New York City's Local Law 144 followed, with enforcement from July 2023: an annual independent bias audit, published — again, a duty of the employer. In January 2023 HireVue engaged DCI Consulting Group to audit its competency-based and game-based algorithms across race, gender and intersectional groups, noting that the ordinance places the obligation on employers and arguing publicly that vendors should share it.
The audits themselves are the richest quantitative record this deployment has, and they are published by its customers. Pfizer's posted HireVue bias audit, summary produced 5 July 2023, states the method: nationwide applicant data from January 2021 to December 2022, pooled across employer implementations, analysed per implementation and then aggregated. On the Communication assessment for intern and new-college-graduate jobs it reports 20,060 male against 9,121 female applicants, selection rates of 0.66 and 0.65 for Top-plus-Middle against Bottom, a gender impact ratio of 0.98, race and ethnicity ratios from 0.87 to 0.96, and intersectional ratios down to 0.82. On the Adaptability assessment for the same population it reports 7,161 male against 3,884 female applicants, a gender ratio of 0.95, and intersectional ratios from 0.79 to 1.01.
In 2025 a peer-reviewed study read all 116 publicly available Local Law 144 bias audits published between July 2023 and November 2024. Gerchick and colleagues found that DCI conducted 20 percent of them; that 54 percent of audits carried at least one impact ratio above 1, the majority of those being DCI audits of HireVue tools deployed by JetBlue, Citizens, Pfizer or Burlington, an artifact of the aggregation and comparator-group method; and that every "silent duplicate" they identified — identical quantitative results republished across or within audit reports — appeared in DCI-conducted audits, all but one describing HireVue tools. The authors could not determine the cause of all of them. These are observations about a measurement regime, not findings of discrimination and not findings of audit fraud. What they show is structural: because the duty binds the employer while the data and the auditor belong to the vendor, the public record filled with vendor-level measurement wearing many client faces.
Only one channel ever attached a price. On 27 January 2022 six Illinois residents filed Deyerler v. HireVue, Inc. in the Northern District of Illinois, alleging that the software collected facial geometry and voice data during virtual job interviews without the disclosures and written consent Illinois' Biometric Information Privacy Act requires. On 26 February 2024 Judge Jeremy C. Daniel granted in part and denied in part the motion to dismiss: claims under sections 15(a), (b) and (d) proceeded; the section 15(c) profit claim was dismissed, on the ground that selling software is not selling biometric identifiers; and the court rejected the argument that the Artificial Intelligence Video Interview Act precludes BIPA claims, holding that the two statutes impose different but concurrent obligations. That last holding is the load-bearing one: the video-interview statute has no teeth of its own, so a ruling that it displaced BIPA would have closed the only enforcement channel that did.
On 25 June 2026 the Circuit Court of Lake County, Illinois granted preliminary approval of a settlement: a fund of $3,750,000 covering an estimated 91,305 people who completed a HireVue interview involving the challenged voice and facial biometrics technology while in Illinois between 27 January 2017 and 25 June 2026, with an estimated $150 per valid claimant subject to pro rata reduction, a claims deadline of 13 October 2026 and a final approval hearing on 28 October 2026. Every word of that has to be carried with its qualifiers. Nothing has been paid. The class was conditionally certified for settlement purposes only. The settlement is expressly not an admission of wrongdoing, and HireVue denies that it collected or possessed biometrics or any other information subject to BIPA. A claim that survives a motion to dismiss is an allegation held plausible. No court and no regulator has ever found this vendor violated the biometric statute, the FTC Act, or any discrimination law.
One more matter is pending and disputed. On 19 March 2025 the ACLU of Colorado, with Public Justice and Eisenberg & Baum, filed charges with the Colorado Civil Rights Division and the EEOC on behalf of a Deaf, Indigenous Intuit employee, alleging that an automated video interview relying on automated speech recognition disadvantaged her and that her request for human-generated captioning was denied before her promotion was refused, citing her communication style. Both companies dispute the charges. HireVue's chief executive Jeremy Friedman called the complaint entirely without merit and stated that Intuit did not use a HireVue AI-based assessment. These are pending allegations, the operator disputes the central factual premise, and nothing in this file rests on them.
Finally, what nobody measures. Independent scholarship bounds the construct rather than the product: Hickman and colleagues, in the Journal of Applied Psychology in 2022, found that automated video-interview personality assessments trained on self-reports showed little evidence of reliability or validity, while models trained on interviewer reports did better with mixed cross-sample reliability, and cautioned vendors and adopting organizations accordingly. Raghavan and colleagues, at FAccT in 2020, found that algorithmic pre-employment vendors' public claims about validation and bias mitigation are largely unverifiable from what vendors disclose. And the structural gap sits underneath all of it: a rejected candidate generates no outcome data anywhere in this loop, so the pooled multi-employer store that builds the models and computes the mandated audits cannot observe the population the models screened out. That is why the audits report selection-rate ratios and never accuracy, and why no error rate for this engine has ever been published by anyone.
The sociotechnical reading
Most cases in this atlas put the instrument and the person who has to live with it inside one organization. This one draws a line through the middle. The party that decides what the engine measures and the party that decides what the measurement does to a candidate are different companies, bound by a purchase order, and every governance event in the record landed on one side or the other and stayed there.
Start with the fact that reads as good news, because it is the most interesting thing here. In about March 2020 the vendor removed visual and facial analysis from new assessment models, and it did so with a number in hand: about 0.25 percent of predictive power in most job models, about 4 percent for customer-facing roles. That is a rare artefact in this whole atlas — an input dropped because its measured contribution had fallen to near nothing while the cost of scrutiny had risen. Read the mechanism rather than the sentiment. The measurement and the removal were held by the same party, which is why the removal was possible at all; and that party's reach was total, which is why it took effect in every one of hundreds of pipelines at once, without any buyer asking. The same concentration that made the good decision cheap would have made a bad one cheap. Nothing in the record required the vendor to tell anyone, and for about ten months it did not.
The audits are where the layer boundary starts doing real damage, and the damage is not the usual one. This is not a deployment nobody looked at. Two audit regimes reached it: a commissioned audit in 2020 and annual statutory bias audits from 2023. What each could see is the point. The commissioned audit covered one representative pre-built early-career use case, examined neither the tool's technical design nor its training data, worked largely through stakeholder interviews, and its report is readable only under a nondisclosure agreement — so the sentence that travelled, that the assessments work as advertised with regard to fairness and bias, is a vendor's summary of a perimeter almost nobody can inspect. The mandated audits publish real numbers with real applicant counts, and they measure selection-rate ratios rather than accuracy, which is a different question from whether the screen is any good. Between the two of them, the record contains no evaluation of what this engine gets right, and a technology-press review of both regimes recorded that neither addressed whether the products improve hiring at all.
Then the transparency regime inverted, and this is the finding worth taking away from the file. Both statutes that reach this deployment place their duties on employers. Illinois requires the employer to notify, explain and obtain consent, and to delete on request — and states no enforcement mechanism. New York City requires the employer to obtain and publish an annual independent bias audit. But the applicant data is pooled across employer implementations in the vendor's store; the auditor is engaged by the vendor; the analysis is run once, over the pool. So a duty aimed at many buyers is discharged with one seller's measurement, and a peer-reviewed study of all 116 public filings duly found identical quantitative results republished under four different employers' names, every one of them in that one auditor's work. Nobody reading a filing can trace its numbers back to the data they were computed on, because that data is not theirs to see — which is exactly why the study's authors could not determine the cause. A disclosure regime that publishes a number nobody can check has produced transparency about a party it does not reach.
Look at what actually moved anything, and the ranking is instructive. An advocacy complaint to a federal consumer regulator produced publicity and no enforcement action of any kind. A commissioned audit produced recommendations and a phrase. A statutory audit mandate produced a public number and no obligation to do anything about it. The one channel that attached a price was a private class action under a state biometric statute that has nothing to do with hiring, and it reached the vendor because a federal court held in February 2024 that the video-interview statute — the one written for exactly this technology — does not displace it. The instrument written for the problem had no teeth; the instrument written for something else did. That settlement, it should be said again, is preliminarily approved rather than paid, carries no admission, and rests on allegations no court has ever adjudicated.
The last thing is the quietest, and it is the reason the board's central pathway runs at zero. A rejected candidate generates no outcome data anywhere. The pooled store that trains the models and computes the audits contains applicants and their scores; it does not contain what happened to the people the screen sent away, because nothing observes them. Everything downstream inherits that. The mandated audits can compute a ratio of selection rates because selection is observable, and cannot compute an error rate because error is not. The vendor's own checking reads the same store and hits the same wall. There is no channel by which a person can contest a score, and even if there were, the record against which the contest would be judged does not exist. What makes this a governance finding rather than a technical limitation is that the loop's blind spot and the loop's transparency obligation point in the same direction: everything the regime requires anyone to publish is a statistic about the people the system kept.
Two boundaries hold, and they are not decoration. Screened people are not modelled: no hiring decision, tier placement, rejection or employment outcome for any person is computed from anything on the diagram, and the impact ratios and class-size figures are recorded external observations from a mandated filing and a court record. And this is the vendor's file. Its client employers run their own deployments with their own evidence, and nothing here asserts anything about any of them — no time-to-hire figure, no diversity figure, no benefit any buyer has reported. The one historical event this file shares with a client-side deployment in the same domain is the removal of the visual channel, told here from the side of the party that removed it.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.