How the safety layer actually works, what it catches, what it does not, and how we tested it. Written to be checked, not to reassure.
Last updated October 9, 2026.
There are two checks. The first is an instant keyword check: a short list of explicit phrases. It runs on every message, including messages turned away by the size limit, the daily cap or an outage, so the crisis lines are still shown on those paths. That list can only ever escalate a message, never downgrade one, so it is a floor underneath the second check, not a replacement for it.
The second is an AI classifier: a separate model call (claude-haiku-4-5) that runs on the messages the AI answers, in parallel with the reply being written, not after it. A message the keyword check has already marked as a crisis skips the classifier and is treated as a crisis directly.
The reply is held back until the classifier returns. If the classifier flags a crisis, whatever was drafted is discarded and a fresh reply is written under an explicit safety directive: name crisis resources for the person's own region in the same message, not a future one, and stay with them rather than just listing numbers and moving on. If the classifier cannot be reached at all, the system fails toward caution and treats the message as a concern rather than assuming it is safe.
Crisis resources are region-aware, not defaulted to one country. The model is told explicitly which country's numbers apply and instructed never to volunteer a number from elsewhere — a number someone cannot dial is worse than none. Where the region is unknown, the reply points to findahelpline.com, which lists free lines by country.
Health and Kind previously ran an older tool (since retired) that relied only on scanning messages for a fixed list of phrases like "suicide" or "kill myself." In testing against realistic ways people actually express crisis, keyword matching alone missed phrasings that carry real risk but use no obvious keyword — "I don't want to be here anymore," "I've written the letters." A keyword list can only ever catch what it was told to look for in advance, which is why the current system uses one only as an instant first check, backed by the model classifier for everything it would miss.
The classifier reads the actual meaning of a message rather than matching fixed phrases, so it can recognize crisis language it has never seen written that exact way before. It is not perfect — see limitations below — but it is a meaningfully different, and meaningfully better, approach than pattern matching.
A real user reported that talking through ordinary worry about a family member's health, an upcoming diagnosis, a loved one's illness, over several messages, could get misread as the user's own crisis. The classifier was erring toward caution on topic heaviness (illness, death) rather than judging whether the person's own message actually showed danger to their own life. The safe default was the right instinct; the missing piece was a middle category between an ordinary hard day and a genuine crisis, not a lower threshold.
We rewrote the classifier's instructions to explicitly separate a topic sounding heavy from a person's own life actually being in danger, added a worked example (a parent, spouse, or child's health scare) as ordinary difficulty rather than crisis, and made explicit that sustained worry about someone else's health across several turns is not itself escalation. We deliberately did not simply lower the threshold — that would have risked missing genuine crisis language instead.
Before shipping, we ran the fix against the real classifier model: every genuine crisis and concern phrasing already in our test set still classified correctly, and the reported scenario, plus new third-party health-worry cases, no longer triggered a false crisis. After deploying, we replicated the original failing conversation live on the production site to confirm it was actually fixed, not just fixed in testing.
We test the classifier against a small set of realistic crisis phrasings — not just the obvious ones, and deliberately including ambiguous, indirect language people actually use, plus messages that sound heavy but are not the user's own crisis. The test is a script kept with our code. In our most recent run (September 29, 2026), against claude-haiku-4-5, of 13 test messages:
From 23 September 2026, when "Share this exchange for review" went live, until early October, nobody could open the review queue. A database access rule left over from an earlier version referred to its own table, so the database refused every read of the reviewer roles and the review page turned everyone away, including the one account that held reviewer rights. Separately, no account held the permission that opens crisis-flagged exchanges, and nothing told anyone when an exchange arrived. Two exchanges were shared in that time, both within twelve minutes of launch, neither crisis-flagged, and neither was read until this was fixed. We found this ourselves, on 1 October, by checking the live system rather than assuming it worked. Nobody reported it.
We corrected the database rule. Who may open the queue, including crisis-flagged exchanges, is now a short named list that does not depend on the database. We added email alerts after an exchange is stored and a scheduled daily count; delivery is not guaranteed. The console shows how many are waiting before anyone opens the text, records who opens an exchange and when, and supports sign-in by an emailed code. The share panel explains the review purpose, who can read it and the retention period, alongside support lines. It does not promise a personal response or round-the-clock monitoring.
Before changing anything we reproduced the refusal against the live database and counted the waiting exchanges from their flags without opening them. The fixed console was run against a production build with automated checks: the sign-in form by emailed code, the count, opening and marking an item, access refused when signed out, and the share panel's crisis lines and wording. An independent review then tried to break the change before it shipped. This entry is here because a page that only ever lists other people's problems is not worth much.
Tapping either button contributes to a daily count: which mode you were in, whether the reply came from a crisis or concern moment, and whether you found it helpful. These feedback records contain no message text, name, account or IP address. Hosting request logs are separate and may contain an IP address, URL and time; see our privacy policy.
This tells us how often, and in which mode, replies are landing well or badly, but on its own it never lets anyone read the actual exchange.
After tapping one of these buttons, you can additionally choose "Share as product feedback" to help us assess and improve replies. You will not receive a personal reply; this does not start a support conversation. It is a separate, explicit choice, off by default: you see the exact text before anything is sent, you can edit or remove any of it, and nothing is sent until you confirm. Crisis-flagged exchanges that a person chooses to share are marked as such, and are read by the same short named review team as every other shared exchange. Each time an exchange is opened, who opened it and when is recorded.
We do not attach your name, account or IP address to a shared exchange, but the text itself may identify you or someone else. Remove identifying details before sharing. The record is not linked to your account, so we cannot reliably locate it using your account alone. It expires 90 days after its first review or 180 days after it was sent, whichever comes first. Expired text is unavailable to reviewers; a daily cleanup job removes the stored record.
Authorized reviewers can use this feedback to decide what needs attention. Marking feedback reviewed does not automatically change an AI reply or fix the reported issue. The system attempts a counts-only email after a submission is stored and schedules a daily count. Email delivery and review times are not guaranteed. This queue is not monitored around the clock and cannot provide personal or urgent support.
Account chat saving is paused. Conversations are saved in your browser only. Every message is also sent to our AI provider, Anthropic, to write the reply and run the safety check; Anthropic's retention is described on the privacy page.
Signing in does not upload your conversations or make them appear on another browser or device. Journals, check-ins, learned memories and safety plans also stay in your browser. You can make an encrypted manual transfer from Memory. We do not store your conversations by default. Messages are processed by our AI provider to write replies, and text you choose to share as product feedback is kept for review.
Account saving previously used readable database storage, not end-to-end encryption; it is now paused. We checked our Supabase project on 7 October 2026: daily backups are retained for seven days, with point-in-time recovery off. Deleted data may remain in those backups until they expire. This is not a promise of immediate removal from every provider copy. See the privacy policy for the full storage and deletion details.
The status page is updated by hand. To report a problem, write to contact@kindnesscommunityfoundation.com.