Skip to content

PAN Lab example

Nevada DETR generative-AI unemployment appeals

The referee who signs: an AI that drafts the ruling

An AI reads an unemployment-appeal hearing and its evidence, then drafts the ruling itself — approve, deny, or modify the claim — along with the written legal decision, and hands it to a referee to sign. Modeled on Nevada DETR's Google-built unemployment-appeals tool. The pitch is speed: a determination that took a referee hours drops to about five minutes. The catch is the review. The whole safety case is that a human signs off, but the tool exists to clear a backlog, and the referee who keeps rejecting the AI is not the careful one — he is the bottleneck. When the model drafts the reasoning and the ruling, and the gate is a self-assessed accuracy score rather than an independent audit, 'a human is in the loop' and 'a human is rubber-stamping' can look exactly the same from the outside.

Stylized model of a documented deploymentPublic benefits & eligibility

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Generative-adjudication class: an AI that drafts the ruling for a referee to sign network: 6 components and 12 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 5 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the generative-adjudication pattern documented in the Nevada DETR unemployment-appeals case file -- not a reconstruction of the actual tool, its model, or the vendor's platform. It is deliberately not a risk-scoring, fraud-flagging, or eligibility-screening model and not an assistive navigation copilot: the system drafts the recommended determination and the written decision themselves, which is what makes this generative adjudication of a due-process entitlement rather than a score handed to a decision-maker.

  • baseline

    The whole safety case is one human referee sign-off, and the deployment's stated purpose is clearing a pandemic-era appeals backlog, so the model-to-referee channel is drawn live and high at baseline: the tool exists to move an adjudication that took as much as several hours down to about five minutes, and critics warn that backlog and speed pressure could hollow out the review -- the referee who consistently rejects the AI is cast as the bottleneck, not the diligent employee.

  • baseline

    The two-worker human loop is drawn as a genuine but thin inhibiting check (present at baseline, not absent): DETR described two state workers involved and a mandatory referee sign-off, so this is not a MiDAS-style removal of the human -- but it is a shallow second look under throughput pressure, and no referee override or rejection rate is published, so the depth of the review is unmeasured. structured-dissent is what turns that sign-off from a courtesy into a scheduled duty.

  • baseline

    The 90 percent accuracy requirement is a self-assessed acceptance threshold graded by state workers on test decisions, not an independent external audit, so the model-side independent check is drawn as an inactive pathway a lever can open; the external retrieval-augmented generation (RAG) legal-research error ranges some analysts cite (17 to 33 percent incorrect, 18 to 63 percent incomplete) come from general studies, not measurements of this system, and the wrong-Nevada-statute and incomplete-document findings from testing were reported by officials as fixed.

  • baseline

    The retrieval corpus includes a database of prior appeals decisions, so the drafter grounds new recommendations on past adjudications -- a decision-record feedback loop that risks entrenching prior patterns, including any historic bias. The record-side reconciliation check is drawn inactive because no reconciliation of a drafted ruling against source Nevada law happens before it joins that corpus, and no independent audit of the feedback dynamics is published.

  • baseline

    The single-vendor cloud egress reflects the documented record: the generative model runs on one dominant vendor's cloud platform, processing sensitive appeal transcripts and evidence (which may contain Social Security numbers, tax, financial and health detail) off-premise, under state-held encryption keys and continental-US data residency but with no claimant consent and no described opt-out. It is drawn privacy-sensitive so the vendor gate has a clear target at the contract; the consent and single-vendor-concentration concerns are institutional, not gauge mechanics.

  • assumed

    Deployed versus planned: as of the most recent available reporting (March 2026) the system was in delayed pre-deployment testing on historical appeals and described as launching in coming weeks, not independently confirmed to be adjudicating live claimant appeals; the earlier 2024 within-months projection had repeatedly slipped. This Lab models the shape the deployment is being built toward, and the repeated slippage is itself a documented fact, not a claim that the tool is operational.

  • assumed

    Appellants are not in the dynamics; the harm mode on this shape is a wrongly-drafted ruling adopted through a deferential sign-off, recorded outside a diagram like this one, never computed here. The bias and hallucination concern a UNLV computer scientist raised is prospective, not a measured demographic disparity, so no differential client harm is estimated, and per project rules no AI model or version identifier is named -- only the vendor platform and the retrieval-augmented approach.

What this example does not show

  • Deployed versus planned: as of the most recent available reporting (March 2026) the system was in delayed pre-deployment testing on historical appeals and described as launching in coming weeks, not independently confirmed to be adjudicating live claimant appeals; the earlier 2024 within-months projection had repeatedly slipped, and that forward-looking window had elapsed by mid-2026 with the status unconfirmed. Treat it as imminent-but-not-confirmed-live, not operational — the repeated slippage is itself the documented fact.
  • The automation-deference or rubber-stamping risk is expert-projected (claimant attorneys, a former U.S. Department of Labor official, and legal analysis) and an opinion-column argument, not a measured outcome: no referee override or rejection rate is published, and the 90 percent figure is a self-assessed acceptance threshold on test decisions, not an independent audit. The external retrieval-augmented generation (RAG) legal-research error ranges some analysts cite (17 to 33 percent incorrect, 18 to 63 percent incomplete) are from general studies, not measurements of this system, and the wrong-statute and incomplete-document findings from testing were reported by officials as fixed.
  • Appellants and the benefits they do or do not ultimately receive are not modeled here; the Lab models institutional propagation only, and the harm mode on this shape is a wrongly-drafted ruling adopted through a deferential sign-off, documented in the case file and measured outside any diagram like this one. The bias and hallucination concern raised by a UNLV computer scientist is prospective, not a measured demographic disparity; no differential client harm is estimated. Per project rules no AI model or version identifier is named — only the vendor platform (Google Vertex AI Studio) and the retrieval-augmented approach.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • Nevada's generative-AI unemployment-appeals tool was justified as a speed measure for a pandemic-era backlog, projecting a drop in referee determination time from as much as several hours to about five minutes per case, with a mandatory human review DETR said adds an estimated 10 to 30 minutes and a required referee sign-off (Director Christopher Sewell said no AI-drafted written decisions issue without human review). Legal scholars, attorneys who represent claimants, and a former U.S. Department of Labor official warned that backlog and speed pressure could hollow out that review and create incentives to rubber-stamp AI outputs -- one attorney noting the time savings only happens if the review is very cursory, and a legal analysis warning staff might feel pressured to authorize AI decisions with haste. That automation-deference risk is expert-projected, not a measured outcome: no referee override or rejection rate has been published, and claimants are not required to consent to AI processing of their appeal.

    empirical
    • Investigative The Markup (Todd Feathers, via Gizmodo), Google's AI Will Help Decide Whether Unemployed Workers Get Benefits (2024) https://gizmodo.com/googles-ai-will-help-decide-whether-unemployed-workers-get-benefits-2000496215
    • Academic Fordham Intellectual Property, Media and Entertainment Law Journal (Dawn Edelman), Speed, Accuracy, and Risk: Nevada's Use of Artificial Intelligence in Unemployment Claims Appeals (2024) http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/
    • Investigative The Nevada Independent (2025, July 22), Nevada will use AI for unemployment appeals; some lawmakers are skeptical (DETR / Google) https://thenevadaindependent.com/article/nevada-will-use-ai-for-unemployment-appeals-some-lawmakers-are-skeptical
  • Nevada's Department of Employment, Training and Rehabilitation contracted Google to build a generative-AI tool on the Vertex AI Studio cloud platform that reads an unemployment-appeal hearing transcript and evidence, retrieves against a corpus of Nevada unemployment law and prior appeals decisions, and drafts a recommended determination (approve, deny, or modify a claim) together with the written decision for a human referee to review and sign. The contract set a 90 percent success requirement self-assessed by state workers on test decisions -- not an independent external audit -- and DETR said it wanted accuracy higher than 90 percent before going live; rollout was repeatedly delayed over less-than-desired accuracy, including the tool citing incorrect Nevada statutes and failing to pull information from all hearing documents, problems officials said were fixed. Reported cost evolved from about 1 million dollars in 2024 to a total of 2.6 million dollars with about 1.1 million spent by early 2026. As of the most recent available reporting (March 2026) the system was in delayed pre-deployment testing on historical appeals and described as launching in coming weeks; it was not independently confirmed to be adjudicating live claimant appeals.

    empirical
    • Investigative The Nevada Independent (2025, July 22), Nevada will use AI for unemployment appeals; some lawmakers are skeptical (DETR / Google) https://thenevadaindependent.com/article/nevada-will-use-ai-for-unemployment-appeals-some-lawmakers-are-skeptical
    • Investigative The Markup (Todd Feathers, via Gizmodo), Google's AI Will Help Decide Whether Unemployed Workers Get Benefits (2024) https://gizmodo.com/googles-ai-will-help-decide-whether-unemployed-workers-get-benefits-2000496215
    • Investigative The Nevada Independent (Eric Neugeboren), Nevada agencies eye artificial intelligence to speed jobless claims, DMV queries (2024) https://thenevadaindependent.com/article/nevada-agencies-eye-artificial-intelligence-to-speed-jobless-claims-dmv-queries
    • Academic Fordham Intellectual Property, Media and Entertainment Law Journal (Dawn Edelman), Speed, Accuracy, and Risk: Nevada's Use of Artificial Intelligence in Unemployment Claims Appeals (2024) http://www.fordhamiplj.org/2024/10/07/speed-accuracy-and-risk-nevadas-use-of-artificial-intelligence-in-unemployment-claims-appeals/

Where this connects

Institutional pressures in this domain

  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.

All of them in context on the Public benefits & eligibility domain page.

Levers available here and the patterns behind them

Documented case histories