PAN Lab example
MyFriendBen benefits screener
One engine under six screeners: a shared benefits substrate
Six state benefit screeners, run by six different nonprofits answering to six different boards, all compute their eligibility math in one open-source codebase owned by a seventh organization. That is the whole board. The tool itself is advisory: it takes an anonymous survey of about six minutes, returns a list of programs a household is likely eligible for with dollar values attached, and submits nothing - every actual determination is made later by a government caseworker who never sees the screen. So watch where an error could travel. It cannot reach a determination, and it cannot reach an official record. It reaches belief: what a navigator working from the report takes to be true, in a job where learning one program's rules the other way takes months. And watch the two checks. The check that would set the encodings against the statutes they compile has never been performed by anyone who did not write them; the only figure in the record that speaks to accuracy comes from the builders, with its method unpublished. The check that would let a wrong estimate come back does not exist, because the screening is anonymous by design - the same choice that makes the intake safe makes the record unusable as memory. In one state, in one year, independent journalism counted about $30M in benefits identified and about $5M estimated to have actually been obtained. Nothing in this system measures that gap. Before you pick a target level: this board cannot be won under Service and Safety Targets or All Governance Targets. The Lab offers this deployment every tool its own record supports, and applying all of them at once costs two to three times the budget and still leaves pathways open, so no amount of money reaches the win condition at those settings. That is a measurement of the deployment this network is derived from, not a puzzle waiting to be cracked. Explore and Service Targets Only can be won.
Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.
What this models
This example runs on the Shared-substrate-class multi-state benefits screener network: 10 components and 20 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.
Evidence base: 5 assumed · 14 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.
- assumed
This models the shared-rules-substrate benefits-screening pattern documented in the MyFriendBen case file, not a reconstruction of the actual deployment. The correlated-failure structure on this board is structural and real: one open-source rules codebase, owned by a separate organization, is the single eligibility-math dependency of six or more nominally independent state screeners. The harm it makes possible is hypothetical. No incident in which an encoding error misestimated benefits across states appears anywhere in the verified record, and nothing here asserts one.
- baseline
The model self-loop is set at the maximum because the sharing is total rather than partial. Every state deployment computes on one codebase and the screener built no rules logic of its own, so an encoding error is perfectly correlated across jurisdictions instead of averaging out, and a correction propagates globally by the same route. That symmetry is the point: the structure that concentrates the risk is the same structure that fixes everything at once.
- baseline
Two input feeds are drawn because two different species of input are documented, and the second one is the case. One is the household's own answers. The other is federal and state statute compiled to computable rules in a codebase owned by a different organization from the one running the screeners. Drawing statute-as-code as a feed rather than as background is what lets the correlated dependency be structure on the board rather than a sentence in the narrative.
- baseline
The check on the encodings is drawn dark, and the reason is a citation gap rather than a refusal. The only figure in the record that speaks to correctness is an accuracy rate above ninety percent, which appears in the builder's, the infrastructure organization's and the state operator's materials with no published methodology and no independent audit. It is treated here strictly as a claim and strictly as an optimistic upper bound. The independent research frame for the whole category is that no systematic evaluation exists of whether benefits screeners increase benefits access.
- baseline
Three operator classes are drawn because the record describes three groups at three different employers with three different kinds of authority. The substrate maintainers, at a separate nonprofit, write the encodings every state depends on. The affiliate and hub product staff configure what appears in a given state, including a materiality cut below which a program is left out entirely. The navigators are employed by independent partner agencies, not by the organization that runs the tool. No one of the three can set the working rules of the other two, and that split is the reason most work-rule levers are not offerable here.
- baseline
Navigator adoption is set at the maximum on a documented mechanism rather than on a usage count. The record gives the reason navigators reach for the tool: training a case manager on a single program's rules otherwise takes months, which is precisely the expertise that would let someone notice a wrong estimate. Weekly use is documented at one employer, where between one hundred and one hundred thirty individuals use it each week, and one partner reports serving about four thousand individuals a year through it.
- baseline
The household inflow is set at the maximum and is deliberately not marked privacy-sensitive, which is the opposite of the usual reading and is derived rather than defaulted. Everything the engine knows about a household comes from a self-reported survey with no name, no identity documents and no agency record to check against, so the inflow is maximal. It is not marked sensitive because identity-free intake is this deployment's strongest documented control: immigration status is optional in one state's deployment and is not collected at all in another, and a screening cannot leak a name that was never asked for. Marking it would misdescribe the design.
- baseline
Exactly one edge is marked privacy-sensitive, and it is the conversation store. The structured questionnaire is built to ask for no name and no documents; a free-text exchange with a conversational assistant can carry whatever a person types, including exactly what the questionnaire declines to ask. The assistant is documented only in the public application repository, is behind feature flags, and its deployment extent across state deployments is not established, so the pathway is drawn light. The compliance assertions attached to the stack, including a health-privacy standard, a security certification and a service-level figure, are unverified first-party claims and are labeled as such rather than relied on.
- baseline
The read back from the store to the front line is drawn dark, and it is the same design fact as the privacy strength. Screening is anonymous, holding no name and no identifier, so there is no prior screen for a navigator to open at the next visit. This is the network's version of the case's central tension: the choice that makes the intake safe is the choice that makes the record unusable as memory, and both effects come from one decision rather than from an oversight anyone skipped.
- assumed
No downstream action system is drawn, and that omission is derived. The tool submits nothing: it routes people to separate government application channels and holds no authority to determine, deny or write into any eligibility system. Every determination is made independently by a government caseworker who never sees the screen. This is what separates the shape from every enforcement-carrying network in the atlas, and it is why the error channel here is belief rather than action.
- assumed
No boundary node and no egress pathway is drawn anywhere on this network. The screening store never syncs with any government system, no third-party disclosure demand or data-sharing arrangement appears in the record, and no breach or paste-out is documented. An undocumented pathway is left off the board rather than added because the shape would look more familiar with one.
- baseline
The reviewer is the commissioned evaluation and the advisory channel, and it is wired in through the record rather than through the model, because that is how it actually reads the deployment: through screening data and through users, not through output. Its interim figures come from 301 users and cover discovery, intent and satisfaction, with the funder's own post describing the evaluation as currently under way rather than complete, and the primary report was not located in open search. Those figures are not an accuracy audit and are not verified enrollment. The outbound advisory pathway is drawn present because the channel has demonstrably changed the product before, when about one hundred families' feedback redirected the roadmap in 2022; the same record states that no oversight body has ever audited or forced correction of the eligibility calculations themselves.
- baseline
The law-feedback edge is drawn low and is exactly as slow as the record says. Aggregate screening data and a calculator derived from the same code, credited with identifying one hundred seventy-seven million dollars, supported advocacy around two state bills, and the builder credits that ecosystem with an eight hundred ten million dollar state child tax credit, which the codebase then had to encode. It is documented in one state, over several years, in builder-published material, so it is drawn as a real loop at low intensity rather than as a routine flow.
- baseline
The screen applied to results is drawn as a bounded automatic check, which is what the documentation describes: every report is framed as an estimate rather than a determination, and an immigration-related question routes to a named legal-aid referral instead of an answer. It is deliberately narrow on the board because it is narrow in the record. The failure the same source calls un-overridable is the household shown nothing, which no disclaimer reaches, and full human discretion sits on the apply side of that report and not on the discourage side.
- baseline
Standing workload is set high and the no-AI counterfactual is set below the competent-human baseline, which is unusual in this domain and is derived from the record rather than borrowed. The workload is not screening throughput, which is self-serve, but the encoding-maintenance load: one codebase holding federal and state statute correct for forty-plus programs in one state and thirty-plus in another, across six-plus jurisdictions whose legislatures change the rules mid-stream, while the deployment adds a state a year and grew from ten to twenty-two languages, on an organization whose scale is indicated by a three hundred thousand dollar federal Phase I award. The counterfactual is set low because the documented alternative is failing at scale: roughly one hundred forty billion dollars in benefits go unclaimed annually by one cited estimate and roughly eighty billion by another, three hundred fifty-one million in unclaimed federal earned-income credit sits in a single state, and one cited projection puts poverty reduction at forty-three percent under full enrollment in seven major programs. A low counterfactual is not evidence that the tool is accurate; nobody has measured that.
- baseline
Impact and scale figures on this board are used only where they bear on structure, because the record's own totals do not agree with each other. Twenty thousand people supported in one release, fifty-five thousand households in a builder case study, one hundred thousand plus households in a later funder restatement, one hundred fifteen thousand plus nationally in a joint release, and sixty-five thousand plus families with over a billion dollars identified on a builder page published in the same period. The denominators shift between identified, applied for, unlocked and delivered. One set of figures comes from independent journalism: in all of 2023 in one state, about five thousand five hundred households were screened, about thirty million dollars in benefits was identified, and about five million dollars was estimated to have actually been obtained. That roughly one-in-six ratio is the record's most load-bearing measured quantity and nothing in the system measures it.
- baseline
Dates and coverage are stated as the verified record has them, not as the roster shorthand had them. The statewide public launch in the origin state was in 2023, after a call-line staff pilot in the autumn of that year; January 2024 is the date of the independent journalism, not of the launch. One further state appears in a single infrastructure-organization post and is absent from the builder's own six-state list, so its live status is unclear in the record rather than merely thinly sourced, and nothing here counts it.
- assumed
The households are not on this network and never enter the dynamics. The people being screened are the served population, and the domain's boundary holds: this Lab models institutional propagation among operators, engines and stores only. The documented user population, with a median annual income just over eight thousand two hundred dollars and a median household size of two, served in up to twenty-two languages, is recorded in the case file. So is the un-overridable failure the record names, the household told it qualifies for nothing that then never applies, which no caseworker, auditor or appeal ever sees because the screening is anonymous and disconnected from agency records. Neither is computed here.
- assumed
Differential effects on households are documented, never computed. A wrong estimate plausibly lands hardest on the households with the least margin, and the served population is documented as low-income and multilingual; but the screening is anonymous by design and the store never syncs with government records, so the evidence base that would support measuring such a disparity does not exist and cannot be built from inside the system. That structural fact is recorded in the case file. This Lab estimates no harm to served people.
What this example does not show
- No outcome for any household is modeled. The Lab reads institutional propagation among operators, engines and stores only; the people being screened are boundary-only. The documented user population, with a median annual income just over $8,200 and a median household size of two, served in up to 22 languages, and the un-overridable failure the record names - an eligible household told it qualifies for nothing that then never applies - live in the case file and are never computed on this diagram.
- The correlated-failure structure is real; the harm it makes possible is hypothetical. One shared rules codebase under six-plus nominally independent state deployments is documented in the technical materials of both organizations involved. No incident in which an encoding error misestimated benefits across states appears anywhere in the record, and nothing here asserts one.
- The accuracy figure is a claim, not a measurement. An accuracy rate above ninety percent appears in the builder's, the infrastructure organization's and a state operator's materials with no published methodology and no independent audit. It is usable only as an optimistic upper bound. Independent category-level research finds that no systematic evaluation exists of whether benefits screeners increase benefits access.
- Nearly every impact total is a self-published or funder claim, and they contradict each other. Twenty thousand people supported, fifty-five thousand households, one hundred thousand-plus households, one hundred fifteen thousand-plus nationally, sixty-five thousand-plus families - alongside dollar totals of $33M, $52M, $58M and $1.2B identified, with the denominators shifting between identified, applied for, unlocked and delivered. Only one set of figures comes from independent journalism: about 5,500 households screened in one state in 2023, about $30M identified, about $5M estimated obtained.
- The commissioned evaluation is ongoing and measures something other than accuracy. Its interim figures from 301 users - a majority discovering a benefit for the first time, most planning to apply within three months, nearly all reporting satisfaction - come from an evaluation the funder's own post describes as currently under way, are cited only through the deployment's own materials, and cover discovery, intent and satisfaction rather than calculation accuracy or verified enrollment. The primary report was not located in open search.
- There is no adversarial primary record. No inspector-general audit, court filing or government evaluation of this tool exists, which is unsurprising for an advisory nonprofit tool outside any statutory scheme. It means harms of the kind this board draws are structurally undocumented, which is not the same as demonstrated absent.
- Dates are stated as the verified record has them. The statewide public launch in the origin state was in 2023, after an autumn 2023 call-line staff pilot; January 2024 is the date of the independent journalism, not of the launch. One further state appears in a single infrastructure-organization post and is absent from the builder's own six-state list, so its live status is unclear in the record rather than merely thinly sourced, and it is not counted here.
- The conversational assistant is drawn light because the record is light. It appears in the public application repository, behind feature flags, proxying a separate service with conversation endpoints. Its deployment extent across state deployments is not established, and no evaluation of it exists.
- Differential effects on households are documented, never computed. A wrong estimate plausibly lands hardest on households with the least margin, but the screening is anonymous by design and the store never syncs with government records, so the evidence base for measuring such a disparity does not exist and cannot be built from inside the system. That is recorded in the case file.
Sources and evidence
What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.
MyFriendBen is an advisory multi-state benefits screener: an anonymous survey of about six minutes returns a report of programs a household is likely eligible for, with estimated dollar values, and submits nothing, routing people instead to separate government application channels that make every determination independently. All of its eligibility math is computed by PolicyEngine, a separate nonprofit whose open-source codebase encodes federal and state statute; MyFriendBen deliberately built no proprietary rules logic, and the same substrate sits under each of its state deployments, so an encoding error would misestimate benefits in every state simultaneously and a correction would propagate to every state simultaneously. That substrate's scale is indicated by a $300,000 NSF POSE Phase I award announced on August 18, 2025. The shared-substrate structure is documented in both organizations' technical materials; no incident of a cross-state encoding error appears in the public record. The accuracy figure attached to it, 'over 90%', is a builder, infrastructure-organization, and operator claim with no published methodology and no independent audit.
empirical- Vendor PolicyEngine, MyFriendBen Launches in North Carolina, Using PolicyEngine API (2025) https://www.policyengine.org/us/research/myfriendben-nc
- Reference GitHub (MyFriendBen org), MyFriendBen/benefits-api repository (2026) https://github.com/MyFriendBen/benefits-api
- Vendor PolicyEngine, National Science Foundation awards PolicyEngine $300,000 grant (2025) https://www.policyengine.org/us/research/nsf-pose-phase-1-grant
- Vendor Code the Dream / MyFriendBen NC (bennc.org), FAQ - MyFriendBen North Carolina (2025) https://bennc.org/faq/
The gap between benefits identified and benefits received is measured only from outside the system. Independent journalism reported that in all of 2023 in Colorado about 5,500 households were screened, about $30 million in benefits was identified, and about $5 million was estimated to have actually been obtained, with the median user reporting an income just over $8,200 a year and a household size of two. The deployment's own impact totals are self-published or funder claims and are mutually inconsistent, ranging across 20,000+ Coloradans, 55,000+ households, 100,000+ households, 115,000+ households nationally, and 65,000+ families, with dollar figures of $33 million, $52 million, $58 million and $1.2 billion identified, on denominators that shift between identified, applied for, unlocked, and delivered. The one commissioned evaluation, by the Urban Institute, is described by a funder's post as currently under way; its interim figures from 301 users cover discovery, intent, and satisfaction rather than calculation accuracy or verified enrollment, and its primary report was not located in open search. The Aspen Institute's Financial Security Program finds that no systematic evaluation exists of whether benefits screeners increase benefits access.
empirical- Investigative The Colorado Sun, New online tool helps Coloradans quickly determine which public benefits they might be eligible for (2024) https://coloradosun.com/2024/01/05/my-friend-ben/
- Vendor Delta Fund, MyFriendBen Is Coming to Washington (2026) https://www.delta-fund.org/myfriendben-is-coming-to-washington/
- Vendor PR Newswire (MyFriendBen/CPAL release), MyFriendBen Partners with Child Poverty Action Lab to Launch Texas Accelerator (2026) https://www.prnewswire.com/news-releases/myfriendben-partners-with-child-poverty-action-lab-to-launch-texas-accelerator-302814420.html
- Academic Aspen Institute Financial Security Program, Bridging the Information Gap: How Organizations Can Assess and Use Benefits Screeners (2024) https://www.aspeninstitute.org/publications/assessing-benefits-screeners/
The deployments are operated by independent local anchors rather than by one organization: Code the Dream with NC 211 in North Carolina, Benefit Illinois's Illinois Benefit Hub, the MASSCAP community-action network in Massachusetts, and Child Poverty Action Lab in Texas, each configuring its own program catalog on a shared white-label backend that carries per-state feature flags and admits a program only if it is worth at least $300 a year. The front-line users are employed by those partner agencies and others: 2-1-1 Colorado piloted the tool with its own call staff in autumn 2023, and 100 to 130 individuals at one Dallas employer use it weekly, with one partner reporting that it serves about 4,000 individuals a year through it. The builder's stated reason for navigator uptake is that training a case manager on a single program's rules otherwise takes months. Screening data has also flowed outward into policy: a Colorado Child Tax Credit Calculator derived from the same code identified $177 million, and the builder credits the surrounding ecosystem with unlocking an $810 million state child tax credit through HB24-1311, a rule the shared substrate then had to encode. The catalog, threshold, uptake, and policy-loop figures are builder, operator, and joint-release claims.
empirical- Vendor Gary Community Ventures, MyFriendBen (case study) (2025) https://garycommunity.org/case-study/myfriendben/
- Vendor PR Newswire (MyFriendBen/CPAL release), MyFriendBen Partners with Child Poverty Action Lab to Launch Texas Accelerator (2026) https://www.prnewswire.com/news-releases/myfriendben-partners-with-child-poverty-action-lab-to-launch-texas-accelerator-302814420.html
- Reference GitHub (MyFriendBen org), MyFriendBen/benefits-api repository (2026) https://github.com/MyFriendBen/benefits-api
- Vendor MASSCAP (Massachusetts Association for Community Action), MyFriendBen (resource page) (2026) https://www.masscap.org/resources/myfriendben/
A single automated rule set applied uniformly and without human review produced tens of thousands of correlated wrongful fraud determinations in the documented Michigan MiDAS case — one flaw repeating at caseload scale rather than averaging out.
empirical- Government Michigan AG, settlement of civil-rights class action (Bauserman, 2022) https://www.michigan.gov/ag/news/press-releases/2022/10/20/som-settlement-of-civil-rights-class-action-alleging-false-accusations-of-unemployment-fraud
- Investigative IEEE Spectrum, Michigan's MiDAS unemployment system: Algorithm alchemy that created lead, not gold https://spectrum.ieee.org/michigans-midas-unemployment-system-algorithm-alchemy-that-created-lead-not-gold
Where this connects
Institutional pressures in this domain
- Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
- Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
- Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
- Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
- Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
All of them in context on the Benefits navigation & public-facing chat domain page.
Levers available here and the patterns behind them
- Check with a second model — Cross-model verification
- Understand the system — Understand the system
- Mark AI-written records — Provenance labeling
- Review on schedule — Oversight cadence & retrospectives
- Review the riskiest first — Risk-tiered oversight
- Store less data — Data minimization
- Upgrade model — Improve the model
Documented case histories
- MyFriendBen benefits screener
- Nava assistive benefits chatbot
- Caddy adviser copilot at Citizens Advice
- GOV.UK Chat
- Mass.gov Virtual Assistant
- Frida (NAV Norway)
- SSA 800-Number Conversational AI Assistant
- EDD Virtual Assistant
- Burokratt
- Singapore's chatbot fleet refresh: eighty scripted engines slated for retirement onto a shared LLM platform
- IRS collection chatbots: expanded and made permanent with no performance measures
- Albert France Services
- Propel in-app SNAP benefits assistant
- GetCalFresh: the nonprofit front door that carried most of California's online SNAP intake
- Benefits Data Trust wind-down