Skip to content

PAN Lab example

pymetrics Soft Skills Platform cooperative audit

The examination that priced its own limits: a paid source-code audit as the deployment

A hiring-assessment company paid a university team to read its source code, and then let them publish what they found. That is the deployment on this board. Not the screening product — the examination of it. In March 2020 the two sides signed a contract that fixed, before any work began, what the examiners could ask, what they would be paid, what they could say afterwards and who would own the silence in between. Four university auditors then spent a summer inside an analysis environment the company provisioned, reading the code that builds a screening model for each client role, running adversarial tests the company never saw, and holding a contractual right to publish a result that made the company look bad. The thing they examined is a real control, and it is the reason this case is in the catalogue at all. Before any screening model may deploy, the pipeline computes its impact ratio on a held-out panel of more than ten thousand players balanced across seven regulatory categories, at both the 50th and the 70th score percentile. If it falls below four-fifths at either threshold, the model does not ship. Nobody tunes it until it passes; the engagement is reworked or abandoned. The examiners found the testing correctly implemented, demographic data absent from the training features with no overt proxies, and their attempts to sneak a biased model past the gate defeated — every path through the code reached the tests. They judged the hundred-item checklist and the mandatory second data scientist a reasonable safeguard against carelessness and against one person acting badly, and they named what remains: two people acting together. Then read the sentence they actually wrote. The company passed the audit, subject to the qualifications and limitations we state. The qualifications were not discovered at the end; they were agreed at the start. The contract put outside the examination the choice of fairness objective, the category set, intersectional groups, back-testing, customised client work, security, privacy — and construct validity, which is the question of whether the games measure anything a job needs. The auditors say plainly that this was beyond their capabilities as computer scientists. The deepest look this pipeline has ever had could not ask what it was for. Everything after that is what happened to the sentence. It was posted, in full, alongside the contract and the budget and the non-compete, which is why a peer-reviewed critique could be written in 2022 arguing that a company that funds an audit and writes its scope is not being audited independently. It also travelled the other way, compressing as it went: an announcement in May 2020 before any result existed, marketing describing an independently audited product, an acquisition release in August 2022 describing an audited platform. The examiners' own word for the arrangement was cooperative — neither internal nor external. Nothing on this board reads a product claim back against the document it came from. And then the channel changed shape. A city mandate arrived in 2023, and the same product is now examined every year by a paid third-party auditor whose summary an employer must publish. It has the continuity the one-off never had. Its 2025 summary reports every calculated ratio clearing four-fifths, the lowest at 0.914 — computed over 283,256 candidates whose gender was known, and excluding 317,513 whose race or ethnicity was not, roughly twice the number it counted. A study of all 116 public audits under that mandate found the same missing-demographic problem across the regime. The deep voluntary look happened once. The shallow mandatory one happens forever. There is no harm finding anywhere in this record. No lawsuit, no regulator, no consent order. What there is instead is a control that worked, examined by people who were paid by the party they examined, on questions that party helped choose, and a verdict that got shorter every time it was repeated. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets, and cost is not what keeps it open. Take every instrument the parties in this record could actually reach, set each one to full strength, and ignore the budget entirely, at a total of fifty-seven against the ten you are given. One pathway is still open at the end. It is the choice of which fifty to a hundred people already doing the job well the model should learn to look for. That choice is not a gap in this deployment's governance. It is the deployment: a screening model is built by pointing at people who already succeeded, and the examination this board is drawn from aimed at exactly that pathway and could not break it. Widening past what anyone in this record holds does not rescue the board either. Take every instrument in the catalogue at full strength, at a hundred against your ten, and the last pathway does close while what the deployment delivers lands short of what the mode asks. That is a measurement of the deployment this network is drawn from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won, and cheaply: two instruments, costing four of your ten.

Stylized model of a documented deploymentHiring & employment screening AI

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Cooperative-audit-class screening pipeline examination network: 13 components and 25 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 2 assumed · 8 published baseline · 4 measured. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • baseline

    The deployment modelled here is the AUDIT ARRANGEMENT rather than the assessment product, following the PAN entry's own scoping statement for this deployment. The screening pipeline is drawn because it is the object examined; every other component on this network is part of the examination, its governing instruments, or what became of its verdict. The assessment product's own client-side story belongs to a different deployment with different evidence.

  • baseline

    No harm finding, litigation, regulator action or enforcement proceeding against this vendor appears anywhere in this record. This network is the catalogue's positive pole: it draws a shipped control that an examination with source-code access found working within a stated scope, and it draws the limits of the channel that examined it. Nothing here should be read as an allegation, and no pathway or pressure on this board rests on one.

  • baseline

    The examination is named cooperative because that is the auditors' own term for it. They place the arrangement outside both the internal and the external categories the field's cited definitions offer: they were not employees, and they held privileged access to source code, data and staff. The phrase independently audited appears in this record only as the vendor's marketing language, which a peer-reviewed critique documents as unsupported by the paper's own cited definitions. This network never uses it in its own voice.

  • measured

    The scope exclusions were pre-agreed with the audited party in the contract before the work began, not discovered as gaps afterwards. Excluded by that agreement: the choice of fairness objective and metric, the regulatory category set, intersectional groups, construct validity of the games, the newer reasoning games, post-deployment back-testing, customised client processes, security and privacy compliance. Whether the games measure anything job-relevant, which is the assessment's reason for existing, was not assessed; the auditors state plainly that it was beyond their capabilities as computer scientists.

  • measured

    The gate is drawn as a bounded automated screen rather than as a second model, following the PAN entry's own reasoning that it does not generate anything, it refuses. Its check pathway is held at the middle rung because the bounds are documented rather than supposed: it tests seven regulatory categories one at a time rather than in combination, against a four-fifths threshold that the peer-reviewed study of the successor mandate describes as a rule of thumb rather than the legal standard, and nothing re-checks adverse impact automatically once a human has signed the model off.

  • measured

    The evaluation-from-outside pathway is drawn at zero on the PAN entry's own recorded field for this deployment, which states that it has had no independent evaluation. That is not a statement that nobody looked. It is a statement about what the looking was: the deepest examination the pipeline ever had was scoped, funded and co-authored by the party that owns it, and the annual channel that succeeded it reports selection ratios without source-code access or adversarial testing.

  • baseline

    PAN carries one public store holding the contract, the protocol, the budget, the report, the paper and the annual mandated summaries, with the compression of the verdict into marketing recorded as an attribute of that store. This network draws three stores from it, because the three are written by different parties under different duties and read by different readers, and because a compression is only visible as a transfer between two records. Every PAN edge touching that store is routed to exactly one of the three; none is duplicated.

  • baseline

    PAN has no edge kind for a check, so two of this network's five checks redraw PAN peer edges: the second reviewer's mandatory read of a completed model, and the audit team's findings reaching the people who authored the engagement. A channel that corrects a model before production is inhibiting in this vocabulary and reinforcing in PAN's; the widths still come from PAN on the stated rung mapping. The other three checks have no PAN counterpart. Nothing is asserted on this side that the PAN entry does not already record.

  • baseline

    This network is drawn at the coarsest granularity at which every documented mechanism of the arrangement is still distinguishable, so several documented flows are carried on a neighbour rather than on a line of their own. Folded this way: the panel results a second reviewer works from, onto the arrival of the completed model; the sign-off recorded against the engagement, onto the second read itself; the auditors reading their own posted terms and the engagement evidence reaching the public record, onto the posting; the supplementary material produced for the auditors, onto their deep read; the annual channel's reading of prior filings, the panel arithmetic it publishes and the statistics it is given, onto the annual filing; the published number returning to the people who build models, onto the filings read back into builds; what the one-time examination passed to the standing one, onto the standing channel; the notebook as a working surface, onto the notebook being authored; and the critique entering the scholarly record, onto the critique itself. Every fact each folded pathway carried is stated on the survivor named here, and no pathway was folded because of a count.

  • baseline

    The scholarly critique channel is drawn on this network although PAN carries no operator class for it, because the record's own account of where independent scrutiny of this arrangement came from is post-publication scholarship: a peer-reviewed paper in 2022 that made this engagement its lead case, and a 2025 study of every public audit under the successor mandate. Their characterisations are the authors' arguments, attributed to them, and adjudicated by no court, regulator or board.

  • baseline

    Screened candidates and assessed players are outside these dynamics. No hiring decision, tier placement, rejection or employment outcome for any person is computed from anything drawn here, and no score over any person is authored anywhere in this bundle. The impact ratios and candidate counts that appear are recorded external observations from a published audit summary, and the demographic counts are that summary's own statement of who its arithmetic could and could not include.

  • assumed

    No error rate, accuracy figure or defect rate for the examined pipeline exists in any source in this record, and its absence is a finding of the case rather than a gap in this model. An impact ratio is a comparison of selection rates between groups; it is not a measure of whether the screen picks the right people, and nothing on this network treats it as one. Every PAN edge for this deployment is marked estimated, and the widths drawn here sit within the rungs those PAN values support.

  • assumed

    The counterfactual floor is authored at a competent human baseline because the record supports both halves of that reading and neither more. The human control is real, mandatory and independently assessed as reasonable against carelessness and one bad actor. The same assessment recorded that no programmatic re-check follows it, and the documented prior state of this market was eighteen assessment vendors whose validation and bias-mitigation claims were largely unverifiable from outside. No measurement anywhere compares this deployment with its own human counterfactual on the same candidates, and nothing here computes one.

  • measured

    The four-fifths gate examined in summer 2020 was a snapshot of the standard, non-customised pipeline. The auditors disclaim any claim about earlier practice, later practice or customised client engagements, and this network carries that limit. The annual summaries published in 2023 and 2025 are the only attestation of current practice, and what they verify is bias-testing assertions and reported statistics, not the continued existence or the exact mechanics of the refuse-rather-than-ship gate.

What this example does not show

  • No litigation, regulator enforcement, or consent decree involving pymetrics, Harver's pymetrics assessments, or the audit was identified as of 2026-08-28; the only legal instruments in the record are the private audit contract and the LL144 compliance regime. Pleadings discipline is moot; the critique literature is scholarship, not proceedings.
  • This network models the audit arrangement, not the assessment product. It shows nothing about how any client employer used the screening tool, what any candidate experienced, or what any hiring outcome was. No score over any person is authored here and no employment outcome is computed from anything drawn.
  • The examination is cooperative, which is the auditors' own term and their own placement of it outside both the internal and the external categories the field's cited definitions offer. The phrase independently audited belongs to the vendor's marketing and is documented by a peer-reviewed critique as unsupported by the paper's own cited definitions. This network never uses it in its own voice.
  • The corporate-capture reading — that four of eight authors were employees of the audited company including its chief executive, that the company funded the work and set the scope, and that a paper negating the company's validation claims would have been unlikely to be submitted — is a peer-reviewed position paper's argument, attributed to its authors. No court, regulator, board or conference process has adjudicated it.
  • The findings are a summer-2020 snapshot of the standard, non-customised pipeline. The auditors disclaim any claim about earlier practice, later practice or customised client engagements, and nothing here extends their verdict to the platform under its later owner.
  • The 2025 mandated summary's ratios are honest arithmetic over a self-selected sub-population and are not evidence of platform-wide compliance. They exclude 317,513 candidates of unknown race or ethnicity, roughly twice the number included in the race calculations. An impact ratio compares selection rates between groups; it is not an error rate, and this network never treats it as one.
  • The peer-reviewed finding of systemic inadequacy across the successor mandate characterises the REGIME — all 116 public audits between July 2023 and November 2024 — and not this product's audits specifically, which were among the more complete in disclosing intersectional tables and unknown-group counts. Nothing here converts a regime-level critique into a finding about these summaries.
  • The payment figures are independent journalism's reporting from the publicly posted budget. The transparency address printed in the 2021 paper no longer resolves; the live page is elsewhere, and this network records the link rot rather than repeating a dead citation.
  • No error rate, accuracy figure or defect rate for the examined pipeline exists in any source in this record, and no measurement anywhere compares this deployment with its own human counterfactual on the same candidates.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • In March 2020 Northeastern University and pymetrics, inc. signed a sponsored research agreement that fixed, before any work began, the scope of a source-code audit of pymetrics' candidate-screening model-generation pipeline, the auditors' remuneration, their publication rights, a non-compete, and a thirty-day responsible-disclosure window. Four Northeastern auditors conducted the audit in summer 2020 inside a pymetrics-provisioned AWS virtual machine, with access to source code, eight Jupyter notebooks, one engagement's full training and evaluation data, confidential technical, fairness-testing, and job-analysis documents, staff time, and a closing live demonstration of a production model trained, tested, and deployed. The independence instruments were real and are documented: pymetrics paid $104,465 to the university, $64,813 of it salaries for the team, structured as a grant and paid in full BEFORE findings were delivered so it could not be conditioned on results, and supplied the audit compute at its own expense; the auditors kept their test methods secret from pymetrics throughout; and they held final editorial discretion including the contractual right to publish results that might reflect negatively on pymetrics. The contract, the audit and data-sharing protocol, the non-compete, the budget spreadsheet, and the final report with section 3.2 redacted for proprietary information were posted publicly by the auditors. The paper was published at ACM FAccT in March 2021 with eight authors, four of them Northeastern auditors and four pymetrics employees including the company's chief executive as last author. The auditors coined the term 'cooperative audit' for the arrangement and explicitly placed it outside the existing internal/external taxonomy: they were not pymetrics employees, yet held privileged access to source code, data, and staff, so by the field's own prior definitions the audit was not an external one. The transparency address printed in the 2021 paper no longer resolves; the live page is at cbw.sh/research/audits/, verified 2026-08-28.

    empirical
    • Peer-reviewed Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
    • Academic Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT '21) https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
    • Academic Wilson, C. Algorithm Audits (transparency page carrying the pymetrics non-compete, sponsored research agreement and work plan, budget, and redacted final report); the URL printed in the 2021 paper is dead and this is the live location as of 2026-08-28 https://cbw.sh/research/audits/
    • Investigative Schellmann, H. (2021, February 11). Auditors are testing hiring algorithms for bias, but there's no easy fix. MIT Technology Review https://www.technologyreview.com/2021/02/11/1017955/auditors-testing-ai-hiring-algorithms-bias-big-questions-remain/
  • The control the audit examined is a pre-deployment four-fifths gate with abandon-if-noncompliant semantics. Per client role, pymetrics built a screening model from an in-group of 50 to 100 high-performing incumbents against an out-group sampled from a player database of more than two million, then tested the candidate model against a held-out 'bias group' — typically more than 10,000 players engineered to equal proportions across the seven EEOC categories, drawn from a pool of more than 600,000 players with self-reported demographics at over 75% survey completion per the vendor. Recommendation tiers sit at the 50th and 70th score percentiles, and deployment requires an impact ratio of at least 0.8 at BOTH thresholds on the bias group. If no compliant model is found, the engagement is reworked or no model ships. The audited code used support-vector-machine models over 64 features produced by 12 core games, filled gaps by median imputation, and dropped players missing more than two games.

    empirical
    • Peer-reviewed Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
    • Academic Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT '21) https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
  • The auditors found pymetrics' code correctly implemented four-fifths testing over the seven EEOC categories; that demographic data was not used as a training feature and no overt proxies such as zip codes were present; that they could not construct manipulated incumbent data producing a biased model without it being flagged, because every control-flow path reached the adverse-impact tests; that the 100-plus-item completion checklist per model plus the mandatory second-data-scientist review before production was a reasonable safeguard against negligence and against a single malicious insider, leaving collusion between two insiders as the residual path they explicitly named; and that median imputation did not substantively alter adverse-impact results despite statistically significant demographic differences in missing data. The auditors also recorded that no programmatic back-end re-check of adverse impact occurs after the data-scientist stage. Their own summary is that pymetrics 'passed the audit, subject to the qualifications and limitations we state'. No litigation, regulator enforcement action, consent decree, or documented harm involving pymetrics or the audit was identified anywhere in this record.

    empirical
    • Peer-reviewed Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
    • Academic Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT '21) https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
  • The audit's exclusions were pre-agreed with the audited party in the contracted scope before the work began, not discovered as gaps afterwards. Excluded by that agreement: the choice of fairness objective and metric, the EEOC category set, intersectional groups, construct validity of the games, the newer reasoning games, post-deployment back-testing, customised client processes, security, and privacy compliance. Whether the games measure anything job-relevant — the assessment's reason for existing — was not assessed; the auditors state it was 'beyond our capabilities' as computer scientists. Intersectionality was excluded because such groups are not EEOC-recognised and all parties agreed the legal risk of optimising to a non-regulatory standard barred it, a reason the record states rather than adjudicates; pymetrics itself expressed interest in it. The audit is a summer-2020 snapshot of the standard, non-customised pipeline, and the auditors disclaim any claim about earlier practice, later practice, or customised client engagements.

    empirical
    • Peer-reviewed Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
    • Academic Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT '21) https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
  • At ACM FAccT 2022, Young, Katell, and Krafft made this engagement the lead case of what they term 'publishing and certifying corporate apologia': four of the paper's eight authors were pymetrics employees including the chief executive as last author; the company funded the work and set the scope, 'precluding various scenarios from scrutiny'; and 'it seems unlikely that this paper would have been submitted had the results negated the firm's validation claims.' They also documented the downstream loop: pymetrics had publicised the engagement in a company post of 13 May 2020, before any results existed, and after publication marketed the product as 'independently audited' — a term the audit paper's own cited definitions do not support, since the auditors classify the arrangement as cooperative and place it outside both the internal and external categories. They further recorded that a co-author sat on the FAccT Executive Committee when the paper was accepted and that 2021 FAccT review did not mandate funding disclosure. These are the critics' characterisations and arguments in a peer-reviewed position paper, attributed to them; no court, regulator, board, or conference process has adjudicated them. The 'audited' framing carried into the acquisition: Harver announced its acquisition of pymetrics on 11 August 2022 with terms undisclosed, describing the product as 'mitigating multiple forms of bias through an audited AI platform'.

    empirical
    • Academic Young, M., Katell, M., & Krafft, P. M. (2022). Confronting Power and Corporate Capture at the FAccT Conference. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT '22) https://facctconference.org/static/pdfs_2022/facct22-3533194.pdf
    • Vendor Harver via PR Newswire (2022, August 11). Harver Acquires pymetrics, Further Enhancing Talent Decision Capabilities Across the Employee Lifecycle https://www.prnewswire.com/news-releases/harver-acquires-pymetrics-further-enhancing-talent-decision-capabilities-across-the-employee-lifecycle-301603823.html
  • The audited practice migrated into a statutory channel. Under New York City Local Law 144, in effect from July 2023, the same game-based assessment platform is audited annually by a paid third-party auditor, BABL AI Inc., and the summary is published by the party deploying the tool: a V2.0 summary dated 29 June 2023, addressed to 'pymetrics inc. (Harver)', was posted by the employer Paramount, and a V1.0 summary dated 17 July 2025 for what is now Harver's Soft Skills Platform was posted by Harver. The platform retains the same three recommendation tiers and a 50th-percentile selection threshold. The 2025 summary reports every calculated impact ratio at or above 0.8, the lowest non-intersectional ratio being 0.914 for Asian candidates, computed across 164,014 male and 119,242 female candidates with known gender. It also reports what the calculation excluded: 191,455 candidates of unknown gender, 317,513 of unknown race or ethnicity — roughly twice the included known-race population — and 331,177 with at least one unknown demographic, with groups under two percent of the sample reported as not applicable. These are recorded external observations from a mandated annual audit: selection-rate ratios over candidates who disclosed demographics, not harm findings, and not evidence of platform-wide compliance. What the annual audits verify is bias-testing assertions and reported statistics, not the continued existence or exact mechanics of the pre-deployment abandon-if-noncompliant gate.

    empirical
    • Government BABL AI Inc. (2023, June 29). Summary of Bias Audit Results: Audit of the pymetrics Soft Skills Platform for New York City's Local Law 144 (V2.0), published by Paramount as a deploying employer https://www.paramount.com/sites/g/files/dxjhpe226/files/2023-07/Harver_Pymetrics-Final_Audit_Summary-2023-06-29.pdf
    • Government BABL AI Inc. (2025, July 17). Summary of Bias Audit Results: Audit of Harver's Soft Skills Platform for New York City's Local Law 144 (V1.0), published by Harver https://harver.com/wp-content/uploads/2025/11/pymetrics-Soft-Skills-Platform-2025-Bias-Audit.pdf
  • An ACLU-led study published at ACM FAccT 2025 analysed all 116 publicly available Local Law 144 bias audits from July 2023 to November 2024 and found the mandated audits to be incomplete evaluations of bias: missing demographic data, opaque aggregation, problematic test data, and metrics that do not represent how the tools are actually used. It warns of audit washing, corporate capture of auditors, and 'discrimination-hacking', records that the four-fifths rule is a rule of thumb rather than the legal standard, and shows that tools reporting four-fifths compliance could be in violation once missing-demographic impacts are considered. It cites the pymetrics cooperative audit among the proposed audit standards the field built on, and the corporate-capture critique among its warnings. The finding characterises the REGIME across all 116 audits and not this product's audits specifically, which were among the more complete in disclosing intersectional tables and unknown-group counts.

    empirical
    • Academic Gerchick, M., Encarnacion, A., Tanigawa-Lau, C., Armstrong, L., Gutierrez, A., & Metaxa, D. (2025). Auditing the Audits: Lessons for Algorithmic Accountability from Local Law 144's Bias Audits. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT '25) https://facctconference.org/static/docs/facct2025-206archivalpdfs/facct2025-final14-acmpaginated.pdf
  • The baseline this engagement stood out against is documented in the prior literature: a survey of 18 algorithmic pre-employment assessment vendors published at ACM FAT* 2020 found their publicly stated validation and bias-mitigation claims largely unverifiable from outside. The pymetrics engagement was the first publicly documented case of an assessment vendor opening its source code to outside auditors under pre-negotiated publication rights.

    empirical
    • Academic Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020). Mitigating Bias in Algorithmic Hiring: Evaluating Claims and Practices. In Proceedings of FAT* '20, 469-481. https://doi.org/10.1145/3351095.3372828 https://arxiv.org/abs/1906.09208
    • Academic Raghavan, M., Barocas, S., Kleinberg, J., & Levy, K. (2020). Mitigating Bias in Algorithmic Hiring: Evaluating Claims and Practices. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAT* 2020) https://arxiv.org/abs/1906.09208
    • Peer-reviewed Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
  • The audited pipeline ran at scale: more than 600 active client engagements between January and October 2020, covering 16 of 23 major O*NET occupation groups, with players from 191 countries and roughly 40% of them in the United States. pymetrics also open-sourced its adverse-impact testing framework as the audit-ai library, cited in the audit paper, which implements the four-fifths rule alongside Fisher's exact test, a z-test, chi-squared, and Cochran-Mantel-Haenszel tests; the repository remains public with low activity after the acquisition. No error rate, accuracy figure, or defect rate for the model-generation pipeline is published in any source in this record, and no measurement compares this deployment with a human counterfactual on the same candidates.

    empirical
    • Peer-reviewed Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. In Proceedings of FAccT '21, 666-677. https://doi.org/10.1145/3442188.3445928 https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
    • Academic Wilson, C., Ghosh, A., Jiang, S., Mislove, A., Baker, L., Szary, J., Trindel, K., & Polli, F. (2021). Building and Auditing Fair Algorithms: A Case Study in Candidate Screening. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT '21) https://www.ccs.neu.edu/home/amislove/publications/Pymetrics-FAccT.pdf
    • Vendor pymetrics, inc. audit-ai (open-source bias-testing library implementing four-fifths, Fisher, z-test, chi-squared and Cochran-Mantel-Haenszel tests), via GitHub https://github.com/pymetrics/audit-ai

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Hiring & employment screening AI domain page.

Levers available here and the patterns behind them

Documented case histories