Skip to content

PAN Lab levers

Lever

Understand the system

The organization funds continuing study of how its AI deployment actually behaves. The people correcting its mistakes learn where the real failures are. They spend their limited time on those failures instead of dividing it evenly across everything.

What it is

Most organizations evaluate a deployment once, before launch, and then assume it behaves as tested. Understanding the system means paying for the work of watching it after launch: sampling its output, tracing its errors to their causes, and telling the people who check it where to look. In the Lab, pulling this lever also makes some other levers on the same network cheaper to pull. Pulling it also lets the levers that depend on understanding work at full strength.

What it pushes on in the Lab

In the Lab, this lever raises the checking and correcting that the people using the system can do. It also weakens the pathway by which people read a contaminated record and believe it, because they know better which sources to check.

increasedPeople or agents using it
dampenedContaminated records read and believed

You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.

The Lab applies this lever to the whole network by design. What it changes belongs to the whole deployment, not to one part you could point at.

Its pattern in the Practice Library

The Practice Library describes the pattern behind this lever:Understand the system

The pressures it answers

These pressures list this lever among the levers that answer them:

A lever answers a pressure when it pushes the other way on something the pressure pushes on.

Where you can pull it

Networks in the Lab that offer this lever:116

Every network that offers it

The evidence behind its effects

The Lab cites these claims from the evidence registry for this lever's effects.

In the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-in disparity from about 20% to 9% relative to the tool acting alone.[4]

goldhaberfiebertprince2019GroundingGovernment evaluationSave

Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

Appears in: PAN framework development

Grounds: capability governance: at-node control; model org: allegheny_afst

Topics: child-welfare

stapletonetal2022AcademicPeer-reviewedSave

Stapleton, L., Lee, M. H., Qing, D., Wright, M., Chouldechova, A., Holstein, K., Wu, Z. S., & Zhu, H. (2022). Imagining new futures beyond predictive systems in child welfare: A qualitative study with impacted stakeholders. 2022 ACM Conference on Fairness Accountability and Transparency, 1162–1177. https://doi.org/10.1145/3531146.3533177

doi.org/10.1145/3531146.3533177

Appears in: PAN framework development; Paramerge authored research

Grounds: deployment audit: Allegheny AFST

Topics: algorithmic-fairness, child-welfare

In the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations erred at about 85% without human review versus 44% with it.[4]

aiincidentdatabaseGroundingInvestigativeSave

AI Incident Database, Incident 373 (MiDAS false fraud claims) https://incidentdatabase.ai/cite/373/

https://incidentdatabase.ai/cite/373/

Grounds: model org: michigan_midas

Pull this lever in the PAN Lab and watch which way it pushes the network.

Open the PAN Lab