Skip to content

PAN Lab levers

Lever

Review the riskiest first

The organization directs its human review to the cases where a mistake would cost the most. Reviewers stop dividing their attention evenly across every case. The checking that remains goes to the cases where an error would do the most harm.

What it is

No organization can check every automated answer with the same care. Risk tiering sorts cases by what is at stake for the person, such as a child's safety, a family's benefits, or a patient's care. The highest-stakes cases go to the most careful review. In the documented cases cited below, human review cut error and disparity roughly in half. It worked because reviewers knew things about the case that the model did not.

What it pushes on in the Lab

In the Lab, this lever raises the checking and correcting that the people using the system can do. It does so by focusing their review on the cases with the most at stake.

increasedPeople or agents using it

You can also pull this lever at a strong tier, which costs more. At the strong tier, each effect below the Lab's strongest setting pushes harder.

In the modes that offer aiming, you can aim this lever at particular parts and pathways of a network. Otherwise it applies to the whole network.

This lever pushes less hard when you switch on lingering effects in the Lab's options, unless you also pull Understand the system.

Its pattern in the Practice Library

The Practice Library describes the pattern behind this lever:Risk-tiered oversight

The pressures it answers

These pressures list this lever among the levers that answer them:

A lever answers a pressure when it pushes the other way on something the pressure pushes on.

Where you can pull it

Networks in the Lab that offer this lever:73

Every network that offers it

The evidence behind its effects

The Lab cites these claims from the evidence registry for this lever's effects.

In the documented AFST evaluation, screener overrides of the tool — roughly a third of its recommendations — cut screen-in disparity from about 20% to 9% relative to the tool acting alone.[4]

goldhaberfiebertprince2019GroundingGovernment evaluationSave

Goldhaber-Fiebert & Prince (Stanford), Impact evaluation summary: Allegheny Family Screening Tool (Allegheny County DHS, April 2019) https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

https://analytics.alleghenycounty.us/wp-content/uploads/2019/05/Impact-Evaluation-Summary-from-16-ACDHS-26_PredictiveRisk_Package_050119_FINAL-5.pdf

Appears in: PAN framework development

Grounds: capability governance: at-node control; model org: allegheny_afst

Topics: child-welfare

stapletonetal2022AcademicPeer-reviewedSave

Stapleton, L., Lee, M. H., Qing, D., Wright, M., Chouldechova, A., Holstein, K., Wu, Z. S., & Zhu, H. (2022). Imagining new futures beyond predictive systems in child welfare: A qualitative study with impacted stakeholders. 2022 ACM Conference on Fairness Accountability and Transparency, 1162–1177. https://doi.org/10.1145/3531146.3533177

doi.org/10.1145/3531146.3533177

Appears in: PAN framework development; Paramerge authored research

Grounds: deployment audit: Allegheny AFST

Topics: algorithmic-fairness, child-welfare

In the documented MiDAS case, error among no-review auto-adjudications ran roughly 93%, and determinations erred at about 85% without human review versus 44% with it.[4]

aiincidentdatabaseGroundingInvestigativeSave

AI Incident Database, Incident 373 (MiDAS false fraud claims) https://incidentdatabase.ai/cite/373/

https://incidentdatabase.ai/cite/373/

Grounds: model org: michigan_midas

In contextual inquiries with Allegheny AFST call screeners, workers calibrated reliance using contextual case knowledge unavailable to the model and reliably detected and overrode erroneous risk scores — complementary human information, not generic distrust, was the safeguard's mechanism.[2]

kawakami2022AcademicSave

Kawakami, A., Sivaraman, V., Cheng, H.-F., Stapleton, L., Cheng, Y., Qing, D., Perer, A., Wu, Z. S., Zhu, H., & Holstein, K. (2022). Improving Human-AI Partnerships in Child Welfare: Understanding Worker Practices, Challenges, and Desires for Algorithmic Decision Support. In CHI Conference on Human Factors in Computing Systems (CHI '22). ACM. https://doi.org/10.1145/3491102.3517439

doi.org/10.1145/3491102.3517439

Appears in: PAN framework development

Topics: algorithmic-fairness, child-welfare, human-ai-interaction

dearteaga2020AcademicSave

De-Arteaga, M., Fogliato, R., & Chouldechova, A. (2020). A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores. In CHI Conference on Human Factors in Computing Systems (CHI 2020). ACM. https://doi.org/10.1145/3313831.3376638

doi.org/10.1145/3313831.3376638

Appears in: Evidence reverification (2026)

Topics: algorithmic-fairness, child-welfare, human-ai-interaction

Pull this lever in the PAN Lab and watch which way it pushes the network.

Open the PAN Lab