Domain Atlas / Hiring & employment screening AI
The 1959 statute and the integrity video screen (Baker v. CVS Health)
Explore this deployment in the PAN Lab ↗
In the PAN Lab, the readouts of this case's model organization carry a shaded evidence band whose width follows the least-established class among the modeling inputs the readings rest on.
The least-established input behind this case's model organization's readings is an assumption, not a measurement. Evidence base: 4 assumed · 8 published baseline.
Around January 2021 a Milton, Massachusetts resident applied for a CVS supply chain position, sat a HireVue video interview and was not hired. As recited in the court's published opinion from the amended complaint, the interview asked integrity-framed questions — what integrity means to the applicant, and a time the applicant acted with integrity — and HireVue uploaded the recordings to Affectiva, a Boston affect-analysis firm spun out of the MIT Media Lab, whose AI 'analyzes candidates' facial expressions, eye contact, voice intonation, and inflection' to draw conclusions about the applicant's degree of cultural fit; HireVue then provided CVS with employability scores. The amended complaint lists the affect features as smiles, surprise, contempt, disgust and smirks, and describes the score as rating traits including conscientiousness and responsibility and an innate sense of integrity and honor. Every mechanical element of that description is an allegation accepted as true for the purpose of a motion to dismiss; none of it was ever adjudicated, and no error rate, accuracy figure or independent evaluation of this screen exists from any source.[3]
What happened
Around January 2021 Brendan Baker, a resident of Milton, Massachusetts, applied for a supply chain position with CVS. He was put through a HireVue video interview and he was not hired. What he was asked, as recited in the court's published opinion from the amended complaint, was what integrity means to him, and to describe a time that he acted with integrity.
What the complaint says happened next is the reason there was a case. HireVue uploaded the recordings to Affectiva, a Boston affect-analysis firm spun out of the MIT Media Lab, whose AI "analyzes candidates' facial expressions, eye contact, voice intonation, and inflection" and draws conclusions about the applicant's degree of cultural fit; HireVue then provided CVS with employability scores. The amended complaint lists the affect features as smiles, surprise, contempt, disgust and smirks, and describes the score as rating traits including conscientiousness and responsibility and an innate sense of integrity and honor. Every sentence of that paragraph is an allegation accepted as true for the purpose of a motion to dismiss. None of it was ever found.
Massachusetts has banned lie detector tests in employment since 1959. In 1985 the legislature added two things to chapter 149, section 19B: a private civil action for any person aggrieved, with a minimum of 500 dollars in damages per violation, treble damages for lost wages or benefits, fee shifting and a three-year limitations period; and a mandatory notice. Every Massachusetts employment application must carry, in clearly legible print, one sentence: "It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability." The provision then sat essentially unenforced for roughly forty years.
The definition is what let a polygraph-era statute reach an affect model, and the drafting choice is precise. The statute covers "any test utilizing a polygraph or any other device, mechanism, instrument or written examination" used "for the purpose of purporting to assist in or enable the detection of deception, the verification of truthfulness, or the rendering of a diagnostic opinion regarding the honesty of an individual." The hook is the purported function. So HireVue's own marketing — quoted in the opinion as claiming capability for "lie detection" and for "screen[ing] out embellishers" — is what brought the assessment inside a definition written for a machine with a cuff and a needle. The vendor's sales copy was the load-bearing evidence, and the vendor's engineering was never examined at all.
Baker filed as a class action. The Boston Globe covered it on 22 May 2023; the federal docket opens on 30 June 2023. He sought to represent all persons who applied for a Massachusetts CVS position, and a subclass of all applicants who sat a CVS HireVue interview alleged to constitute a lie detector test, seeking the statutory 500 dollars per violation plus fees. The defendants were CVS Health Corporation and CVS Pharmacy, Inc. HireVue and Affectiva were not sued. The party that built the inference and the party that ran it both sat outside the case, and the party that bore the duty was two organizational hops from the judgement being made.
On 16 February 2024 Judge Patti B. Saris denied both of CVS's motions in full — the failure-to-state-a-claim motion aimed at the notice count, and the separate motion challenging Article III standing. No count and no defendant was dismissed. On standing the court held that denial of information to which a plaintiff has a legal right can be a concrete injury in fact, and that the notice "would have specifically informed Baker that the HireVue Interview was a lie detector test." That is the mechanism of the injury, stated exactly: the missing sentence removed the option to refuse before pressing record. On the merits count, CVS did not challenge the sufficiency of the allegations that the screen itself violated the lie-detector prohibition, so the court accepted that characterization as plausibly pleaded rather than deciding it.
There is a timing tension in this record and it belongs in the open. HireVue removed facial analysis from new assessments in March 2020 and announced the change in January 2021 alongside an audit it had commissioned; it retained speech-and-language analysis. It told the Boston Globe that "visual and audio analysis have since been eliminated," and its chief data scientist rejected the deception-detection characterization outright, saying the assessments measure work competencies "statistically linked" to job success using "validated industrial organizational psychology." Baker applied around January 2021. Whether Affectiva affect analysis in fact ran on his interview, or on any class member's, was never adjudicated. This file does not assert that it did.
Nor was the scientific question reached. Leonard Saxe, the Brandeis psychologist whose work on polygraph validity helped motivate statutes like this one, told the Globe that there is no neurological signal of deception and "no way for an automated system to distinguish a falsehood from the truth." That is the critique the 1959 legislature was answering, transposed onto a newer instrument. It is recorded in a newspaper, not in a finding.
The case ended without deciding anything. A settlement notice was filed on 17 July 2024; the docket shows the case terminated on 22 July 2024; a stipulation of voluntary dismissal with prejudice followed on 20 September 2024. The settlement was individual and confidential, reached before any class-certification ruling. No monetary terms, no practice changes and no admission of liability were disclosed. The named claim was extinguished and the class went unrepresented.
What propagated instead was the pleading. In roughly a year after the ruling, more than twenty class actions were filed under section 19B, largely by a single New York firm, often with the same individuals suing multiple employers — including Procter & Gamble. Most allege only that an application lacked the statutory notice, with no screening device of any kind involved. Compliance advisories now tell every Massachusetts employer to print the sentence, in the application itself rather than in a policy filed elsewhere, and illustrate the exposure with arithmetic: two hundred applications a year at five hundred dollars is roughly a hundred thousand dollars. That figure is an illustration of the statutory structure, not a measured exposure for CVS or anyone else.
So the correction that is actually spreading across Massachusetts hiring is a sentence on a form, driven by per-application statutory damages and copycat litigation rather than by any adjudicated finding about any technology. It restores something real: an applicant's chance to know what an instrument claims to be before agreeing to sit it. It touches nothing about how a score is produced, by whom, or whether the inference behind it means anything — and on that question this record contains no measurement at all.
The sociotechnical reading
Read this deployment as a duty that sits in one place and a judgement that is made in another. The employer holds the statutory obligation; the platform captures the interview and composes the score; the affect inference, as pleaded, runs at a further company again. Two organizational hops separate the party that answers from the party that infers, and the case's whole shape follows from that gap. The employer could describe what it had bought but not what it did. The companies that could describe it were never asked, because they were never parties, and neither one's internal pipeline was ever put in evidence. The result is that the authoritative public account of this system is a court's recitation of one side's pleadings — a description accepted as true for the purpose of a motion, which nobody has ever reconciled against the thing it describes.
The failure that was litigated is not a model failure. It is an information channel that did not run. The legislature's design was straightforward: tell every applicant, on the form, that this kind of instrument is unlawful, and the applicant can decline before anything is recorded. That one sentence is the entire applicant-facing governance in the record — there was no opt-out, no view of the affect analysis or the employability score, and no channel to challenge an assessment. The court found the missing information itself to be the injury, which is why the accuracy of the screen never had to be litigated at all. A screen can be perfectly accurate and still owe someone the chance to refuse it.
The statute reaches the technology through what the technology was sold as doing rather than through what it does. That inverts the usual direction of AI governance, where a claim about capability is a marketing question and a measurement of capability is the regulatory one. Here the marketing copy is the regulatory fact: an instrument used for the purpose of purporting to detect deception is inside the definition, so a vendor claiming lie detection wrote the evidence against its buyer. It also means the statute bites hardest where a vendor overstates, and it does not bite at all on a vendor that makes the same inferences and describes them cautiously.
The pressure the network runs under is a control that stayed frozen while its subject changed. The notice requirement rendered perfectly for forty years, legible on every form it was printed on and unenforced everywhere it was absent, because enforcement of this statute is a private civil action and nothing else — no regulator investigated this deployment, no agency acted, and no audit of the screen or of the vendor chain exists in the record. A dormant control does not raise an alarm. It waits for someone with standing to read it, and what woke this one was a technology the drafters could not have described.
Then look at how the correction moved, because it did not move through the case. The case settled and produced nothing. What moved was the published ruling: a pleading-stage holding, read by law firms, turned into advisories, turned into twenty-plus follow-on suits, turned into lines appearing on employment applications across the state. The mechanism is per-application statutory damages, which make a form's contents worth checking at scale — and the suits that followed mostly allege no screening device at all. The system-level improvement here is real, and it was produced by litigation risk rather than by anyone learning anything about the instrument. Nothing in this record measures the screen, and nothing in this record ever will: the party with the price attached to it settled, and the parties with the data were never in the room.
The concepts used in this reading are defined in the Field Guide; the governance responses live in the Practice Library. The model organization for this case can be stress-tested in the PAN Lab.