Stillwater
A drop-in middleware that sits between any LLM and a mental-health app's chat UI, monitoring each turn for escalating risk and triggering clinician handoff or crisis-line routing when thresholds are crossed.
Product and clinical leads at digital mental-health startups deploying AI companions
- Multi-turn risk classifier with calibrated sensitivity per conversation length
- Protocol-guided response rewrites that preserve rapport before redirecting to human care
- Clinician dashboard showing live conversations flagged for review with rationale
- Audit-ready logs aligned with HIPAA and emerging FDA software-as-a-medical-device guidance
The cited paper shows risk-detection lifts clinician-preferred escalation by 25-60pp; regulators and insurers are pressuring AI-mental-health vendors right now after a series of harmful incidents.
Multiple recent academic and regulatory signals show urgent need: medRxiv clinician-labeled study on LLM crisis detection limits, FDA Digital Health Advisory Committee met Nov 2025 specifically on gen-AI mental health devices, STAT reports FDA confronting therapy chatbot risks, and public research artifacts like CrisisGuard and MindMate confirm active developer interest.FDA digital advisers to confront risks of therapy chatbots - STAT ↗Suicide- and crisis-risk detection using large language models in mental health chatbots ↗
General-purpose LLM guardrails exist (Guardrails AI, Galileo, Lakera, NeMo Guardrails) but none are purpose-built for multi-turn mental-health escalation with clinician handoff; however, incumbents could expand down-market, and failed prior entrants (Kintsugi closed Feb 2026, Mindstrong previously shut down) suggest the vertical is hard.Mental Health Voice Biomarker Kintsugi Closes, Makes All Technology and Research Public ↗5 Best AI Guardrails Platforms Compared in 2026 | Galileo ↗
B2B middleware can charge per-API-call or per-active-user like Guardrails/Galileo, and regulatory/insurer pressure should support willingness to pay; however, digital mental-health startups are notoriously capital-constrained and adjacent behavioral-health AI vendors (Kintsugi, Mindstrong) have already failed, capping realistic ACVs.Guardrails AI Review (2026) - MakerStack ↗Mental Health Voice Biomarker Kintsugi Closes, Makes All Technology and Research Public ↗
FDA scrutiny, ongoing harmful-incident reporting, and potential future mandates create a long regulatory tail; once integrated as middleware, switching costs are high and crisis-line routing infrastructure becomes durable utility-layer software.U.S. FDA and CMS Actions on Generative AI-Enabled Mental Health Devices Yield Insights ↗Between Help and Harm: An Evaluation of Mental Health Crisis Handling in LLMs ↗
Technically achievable — recent papers show measurable detection lift and open-source crisis-detection frameworks already exist — but requires clinical validation, HIPAA/SOC2 compliance, and partnerships with crisis lines (e.g., 988), making it non-trivial to ship at production grade.CrisisGuard: Crisis Detection System for Mental Health LLMs ↗Domain-Specific Constitutional AI: Enhancing Safety in LLM-Powered Mental Health Apps ↗