Skip to content

PAN Lab example

Calgary Drop-In Centre

The canvas rather than the answer: interpretable screening a shelter's own staff choose to check

A shelter and a university build a deliberately simple screening tool — explicit rules like "81 or more stays in ninety days" that flag chronic shelter use months earlier than the official definitions — and then, instead of handing staff a score, they build an interface that shows the raw client history and lets the staff read it themselves. Modeled on the Calgary Drop-In Centre's interpretable shelter-use screening and its frontline data-navigation interface — their shape, not the real tools. Then they studied their own staff for two and a half years. The finding: the staff deferred to the data by the stakes. For a high-stakes barring decision they read everything and treated the data as "the canvas" they paint on; for lower-stakes housing triage they were more willing to let the data recommend. The tool scores no one and denies no one. So the real question is not whether the model is accurate. It is whether the low-stakes end of that deference holds, and whether anyone ever checks the records the staff themselves author — because the client who quietly stops showing up in the log is invisible to the very evidence meant to surface them.

Stylized model of a documented deploymentHousing & homelessness services

Open this example in PAN Lab v0.1 to apply pressures and levers and watch what the system does.

What this models

This example runs on the Calgary-Drop-In-class interpretable shelter-use screening & frontline data-navigation layer network: 5 components and 13 pathways between them. Every context in the Lab is a stylized model, never a reconstruction of any actual deployment, and each assumption behind it carries a provenance label.

Evidence base: 3 assumed · 3 published baseline. In the Lab, the shaded evidence band behind each headline readout draws its width from the least-established class below.

  • assumed

    This models the interpretability-first, directly-observed-deference pattern documented in the Calgary Drop-In Centre shelter-ML case file — not a reconstruction of the actual tools. The atlas-relevant object is the topology: deliberately interpretable screening (explicit stay-count thresholds and database-queryable rules derived from the shelter's own administrative records) surfaced to frontline NGO staff through a co-designed data-navigation interface that presents raw client histories rather than a risk score, built by a University of Calgary research group in partnership with the shelter. The deployed, studied artifact is the data-navigation interface, not an automated threshold screener; the model scores no one and writes nothing to the record. Any reading that implies an opaque risk score or an automated determination misreads it.

  • baseline

    The defining property is that the model-to-operator deference channel here is directly observed rather than assumed. Across a 2022 to 2024 embedded deployment study (16 staff across 7 role categories; 29.5 hours of qualitative data; five committee observations; three co-design sessions; three deployed versions), the research team documented a stakes-dependent 'data-outsourcing continuum': staff were reluctant to outsource high-stakes barring decisions - treating the data as a starting point for collaborative discussion and as the canvas they then paint on - while reporting more willingness to accept automated data-driven recommendations for lower-stakes housing triage. The two operator classes read the same evidence layer and defer to it differently, calibrated to the stakes; that difference is the payload.

  • baseline

    Underneath the deference channel sits a closed recording loop. There is no machine write to the record - the interface surfaces, humans author - so the loop runs operator to store to model to operator: barring, counselling, and log entries staff author become the same discretionary records the screening and interface read back. The consequence is documented: the clients who 'fly under the radar' have many sleep entries but few bar or counselling events, so the record's sparsity for them is itself a product of staff recording behavior, and a raw history read as complete quietly drops them. The routine attendance store is objective and high-volume; the discretionary log store is sparse and staff-shaped, and it is where the omission concentrates.

  • baseline

    The two load-bearing absences are drawn as inactive checks. First, the record-against-record reconciliation: no ground-truth audit reconciles the discretionary barring/log record against the routine attendance record to surface whom the log systematically omits - the under-the-radar cohort is visible in the attendance counts but invisible in the discretionary record, and nothing routinely reads one against the other. Second, the independent peer evaluation: no independent evaluation of the deployed interface exists - all deployment evidence is authored by the embedded participant-research team, with no usage logs, override or agreement rates, or decision volumes, so nothing outside the team checks whether the low-stakes deference is calibrated or whether the tool helped. These are the case's distinctive safety shape: a tool deliberately built to be audited by its users still needs a check its users cannot supply.

  • assumed

    The evidence for the tools' effect is thinner than the evidence for their design, and its provenance is kept explicit. The model papers are peer-reviewed and quantitative (dataset sizes, threshold definitions, identification-time comparisons, an example rule at about 85% recall / 60% precision), but the deployment study that carries the deference finding is a participant-researcher-authored preprint with no independent venue or DOI found as of mid-2026, and no usage logs, override rates, or outcome data for the deployed interface are published. Model performance figures are research metrics on historical data, not operational reliability claims, and no fetched source confirms the thresholds running as an automated production screener - the deployment is characterized as staged adoption of a data-navigation aid.

  • assumed

    Served people — adults experiencing homelessness, and the shelter, barring, and housing outcomes they do or do not receive — are not in the dynamics; this Lab reads institutional propagation only. No client outcome, override rate, or demographic disparity is computed here, and the fetched record contains no demographic-bias or equity audit of the screening or the interface. The harm surface the model reads is institutional: an omission in the discretionary record, a stakes-dependent deference gradient, and a recording loop that treats staff-authored records as ground truth. Counts vary across the source papers by inclusion window (34,577 vs 41,935 clients in the research datasets; 500+ beds; nearly 7,000 unique clients a year) and should be cited per-paper, never merged. A stay, a bar, or a log entry on this map is an institutional signal, never a person.

What this example does not show

  • Served people — adults experiencing homelessness, and the shelter, barring, and housing outcomes they do or do not receive — are not modeled here; the Lab reads institutional propagation only. This tool scores no one and makes no determination: it surfaces raw client histories to frontline staff who decide. No client outcome, override rate, or demographic disparity is computed here, and the harm surface is institutional — an omission in the discretionary record, a stakes-dependent deference gradient, and a recording loop read back as ground truth — never an individual determination.
  • The threshold and rule figures (81-plus stays in 90 days; a median of about 98 days versus 285 under the Government of Canada definition and 365 under the Alberta definition, roughly 190 to 270 days earlier; an example rule at about 85% recall and 60% precision) are peer-reviewed research metrics on historical shelter data, not measures of operational reliability, and no fetched source confirms the thresholds running as an automated production screener — the deployed, studied artifact is the raw-history data-navigation interface. The deference finding comes from a single participant-researcher-authored deployment study (a preprint with no independent venue or DOI found as of mid-2026); no usage logs, override or agreement rates, or outcome data for the interface are published. A safe-looking baseline is a property of this model, not a safety promise for any real deployment.
  • The staff were reluctant to outsource high-stakes barring decisions and more willing to accept recommendations for lower-stakes housing triage — a stakes-dependent gradient, reported as the staff's own articulated practice, not a measured override or agreement rate; 'reluctant' and 'more willing' are the honest words, not 'refused' or 'trusted.' Counts vary across the source papers by inclusion window (34,577 versus 41,935 unique clients in the research datasets; 500-plus beds; nearly 7,000 unique clients a year, 6,839 in 2022-23) and are cited per-paper here, never merged; the shelter is described in the sources as the largest emergency shelter in Calgary, a Calgary-scope superlative and not a national one.
  • The fetched record contains no demographic-bias or equity audit of the screening or the interface. The papers characterize under-the-radar clients missed by standard definitions — a related but distinct analysis — but no subpopulation error rates are published, so none are asserted or estimated here. This is a decision-support and data-navigation system built by a University of Calgary research group with an NGO shelter operator; its atlas relevance is the directly observed, stakes-dependent deference gradient and the closed recording loop, and any reading that implies individual scoring, automated determination, or a municipal risk-scoring tool misreads it.

Sources and evidence

What this example rests on, claim by claim. Every entry resolves to the same ledger the Evidence Registry publishes.

  • At the Calgary Drop-In Centre, a University of Calgary engineering group and the NGO shelter operator built deliberately interpretable screening for chronic and episodic shelter use - explicit stay-count thresholds (for example 81 or more stays in a 90-day window) and database-queryable rules derived from the shelter's own administrative records, reported to flag candidate clients at a median of about 98 days versus 285 days under the Government of Canada definition and 365 under the Alberta definition - and, rather than surface a risk score, deployed a co-designed data-navigation interface that shows frontline staff raw client histories; no fetched source confirms the thresholds running as an automated production screener, and the deployed, studied artifact is the raw-history interface.

    empirical
    • Academic Messier, Tutty, John, The Best Thresholds for Rapid Identification of Episodic and Chronic Homeless Shelter Use (arXiv:2105.01042 full text, 2021, v3 2023) https://arxiv.org/abs/2105.01042
    • Academic A Rule Search Framework for the Early Identification of Chronic Emergency Homeless Shelter Clients (arXiv:2205.09883, 2022, v3 2023) https://arxiv.org/abs/2205.09883
    • Academic Masrani, Messier, Voida, Dimitropoulos, He, Understanding Data Usage when Making High-Stakes Frontline Decisions in Homelessness Services (arXiv:2510.14141, 2025) https://arxiv.org/abs/2510.14141
  • Across a 2022 to 2024 embedded deployment study of the interface (16 staff across 7 role categories; 29.5 hours of qualitative data; five committee observations; three deployed versions), the participant-research team documented a stakes-dependent 'data-outsourcing continuum': staff were reluctant to outsource high-stakes barring decisions, treating the data as a starting point for collaborative discussion, while reporting more willingness to accept automated data-driven recommendations for lower-stakes housing triage; the finding is the staff's own articulated practice rather than a measured override or agreement rate, all deployment evidence is authored by the embedded research team, and no independent evaluation, usage logs, or decision volumes are published.

    empirical
    • Academic Masrani, Messier, Voida, Dimitropoulos, He, Understanding Data Usage when Making High-Stakes Frontline Decisions in Homelessness Services (arXiv:2510.14141, 2025) https://arxiv.org/abs/2510.14141
    • Academic The Human Behind the Data: Reflections from an Ongoing Co-Design and Deployment of a Data-Navigation Interface for Front-Line Emergency Housing Shelter Staff (CHI 2023 Extended Abstracts, ACM, pp. 1-7) https://arxiv.org/abs/2310.13795

Where this connects

Institutional pressures in this domain

  • Workload surge — Demand outruns staffing; per-case attention shrinks and review becomes triage.
  • Austerity & recovery incentives — Cost-cutting and overpayment-recovery targets tilt the system toward denial and enforcement errors.
  • Vendor opacity — The deploying institution cannot inspect the model, data, or update pipeline it is accountable for.
  • Data & policy drift — The world, the intake process, and the rules change under a system trained on how things used to be — two mechanisms with different remedies: the statistical properties of what the system processes move (concept drift), or the mixture of inputs arriving in deployment differs from the mixture it was trained on (covariate shift).
  • Compliance over substance — Paper controls (sign-offs, checklists) satisfy audits while the behavior they describe erodes.

All of them in context on the Housing & homelessness services domain page.

Levers available here and the patterns behind them

Documented case histories