Skip to content

PAN Lab example

LyssnCrisis counselor QA at ProtoCall Services (988)

The AI watches the counselor rather than the caller: a crisis-line QA scorer

For once, the AI is pointed the safe way round. It never answers a caller in crisis. It listens to the counselor's own call and scores their practice — did they check for suicide risk, did they stay present — then hands the counselor and their supervisor a fidelity dashboard within minutes. Modeled on LyssnCrisis at ProtoCall Services. It closes a real gap: hand review reached under 3% of calls, and now nearly all of them get measured. The catch is second-order. The tool watches the counselor, but nobody independent watches the tool: the trial built to check whether its scores are right, and whether they actually help, has not published. So the question is not whether a caller gets a wrong answer. It is whether a measurement no one has calibrated quietly becomes the yardstick a counselor's career is judged by.

Stylized model of a documented deploymentBehavioral-health & crisis triage

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Lyssn-class counselor-QA scorer on a crisis line network: 5 components and 11 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 2 assumed · 4 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the counselor-quality-assurance pattern documented in the LyssnCrisis at ProtoCall Services case file — not a reconstruction of the actual tool. The defining choice, kept explicit throughout, is direction: the AI scores the counselor's own practice on the supervision link over the operator network, and never acts on a caller. It is a deliberate contrast to the atlas's caller-facing crisis cells (a caller-queue severity ranker and an e-triage intake chatbot), which place the model on the client-facing link.

  • baseline

    The load-bearing pathways are the two model-to-operator feedback edges — the fidelity result fed back to the counselor (shaping future calls) and to the supervisor (coaching and performance management) — plus the store those scores train on. Structurally this is a measurement and state-feedback channel on operator behavior, so the governing levers protect the human read and the calibration of the measurement, not the model's raw accuracy.

  • baseline

    The tool attacks a directly measured coverage constraint: the mandated manual quality assurance (QA) reviewed under 3% of calls, mostly at random, against a guideline that a risk assessment should occur in every call. The manual review is drawn as a real but thin inhibiting check (small random sample), and the tool expands measured coverage toward the full call stream rather than replacing a working full process.

  • baseline

    Peer pathways are authored on both signs: coaching-to-the-score habits and the documented burnout drift away from harder techniques spread across the counselor pool (reinforcing), while a thin collective flag-each-other habit and the mandated manual review run as inhibiting checks. The reliability evidence is real but vendor-authored (four equity-holder authors, three cofounders) with no independent replication.

  • baseline

    The defining latent absence is a dormant independent-calibration check on the fidelity scores: at baseline no non-vendor evaluation validates the scores against ground truth. The registered randomized crossover trial (81 call-takers) was built to test exactly this loop's gain, but its results are unpublished as of mid-2026 and record-level data are proprietary, so every counselor-skill-improvement claim remains a vendor claim. The check is drawn as a latent absence rather than an open pathway, and the forward governance that would stand behind the scores — an independent conformity assessment, a standing oversight cadence, a vendor gate securing independent test access — is what the trial's unpublished result leaves missing.

  • assumed

    Served callers and the crisis counselors themselves are not in the dynamics, and no suicide, crisis, or clinical outcome is computed here — a mandatory boundary for a behavioral-health cell. The harm mode on this positive pole is second-order and never touches a caller directly: a miscalibrated fidelity score contaminating the supervision memory the tool learns from, or score-driven performance management running ahead of the unpublished effectiveness evidence. It is recorded outside a diagram like this one, never a punitive flag on a person.

What this example does not show

  • Served callers and the crisis counselors themselves are not modeled here, and no suicide, crisis, or clinical outcome is computed from anything in this diagram; the Lab reads institutional propagation only, and any such outcome is documented in the case file and measured outside a diagram like this one.
  • The direction is the whole point and must not blur: the tool scores the counselor's practice and never interacts with a caller or issues any caller-facing decision, so nothing here models triage of callers or clinical care.
  • Every counselor-skill-improvement claim is a vendor claim: the reliability study is vendor-authored (four equity-holder authors, three cofounders) with no independent replication, and the registered randomized trial that would test whether the feedback actually helps completed in late 2025 but has published no results as of mid-2026, with record-level data marked proprietary.
  • A safe starting baseline is a property of this model, not a safety promise for any real deployment; the reliability figures are agreement statistics against human raters, not an absolute error rate, and the framing of the scores as a way to identify under-performing staff is aspirational commentary in the case record, not an evaluated practice.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • An AI quality-assurance tool deployed on a national 988 backup line scores crisis counselors' own call practice rather than callers, expanding measured review from the under-3% of calls that had been reviewed by hand toward nearly all of them; a peer-reviewed reliability study of 476 labeled calls reported agreement with human ratings at 98 percent of human interrater agreement for detecting any risk assessment, with average F1 of about 0.86 at call level and 0.66 at statement level, and its authors include four holders of equity in the vendor.

    empirical
    • Academic Imel, Pace, Pendergraft, Pruett, Tanana, Soma, Comtois, Atkins, Machine Learning-Based Evaluation of Suicide Risk Assessment in Crisis Counseling Calls (Psychiatric Services, 2024;75(11):1068-1074) https://pubmed.ncbi.nlm.nih.gov/39026467/
    • Investigative Aguilar, A 988 operator faced with a flood of calls turns to AI to boost counselor skills (STAT News, 2023) https://www.statnews.com/2023/06/22/988-suicide-hotline-lyssn-protocall-artificial-intelligence/
    • Government NIH RePORTER (National Institutes of Health), Voice-based AI to scale evaluation of crisis counseling in 988 rollout (R44MH133517) (2025) https://reporter.nih.gov/project-details/10983779
  • The registered randomized crossover trial of the tool's counselor feedback (81 call-takers) completed on October 31, 2025, but as of mid-2026 no results were posted to the trial registry or found in the peer-reviewed literature and participant-level data were marked unavailable for proprietary reasons, so reported counselor-skill-improvement effects remain vendor claims pending independent publication.

    empirical
    • Government ClinicalTrials.gov (U.S. National Library of Medicine), Voice-Based AI to Scale Evaluation of Crisis Counseling in 988 Rollout (NCT06299384) (2026) https://clinicaltrials.gov/study/NCT06299384
    • Vendor Lyssn.io, Academic Papers: Deployment and evaluation of Lyssn's risk and safety assessment tool at a national crisis and 988 call center (research index, 2026) https://www.lyssn.io/resources/academic-papers/

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Reviewer bottleneck — One fixed-capacity checking stage sits between AI output and consequence; everything queues behind it.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Deadline pressure — Statutory or managerial timeliness rules reward fast approval of machine output over slow disagreement.
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Behavioral-health & crisis triage domain page.

Levers available here and the patterns behind them

Documented case histories