Anywhere a generative system talks to a real person over time.
Not only mental health apps. Generative AI dropped into slots where scripted bots used to sit and nothing else changed — service, benefits, HR, tutoring, clinical intake. The label on the org chart never changed. What is doing the talking did.
The moment an app invites confidences, gives advice and keeps the conversation going, risk arrives automatically. Not because anyone meant harm, but because nobody is holding it to a standard when the conversation drifts. And a content filter cannot see drift, because a content filter reads one message at a time.
What holds a person to this standard, and what holds an AI.
Not only in mental health apps. In tutoring bots, workplace assistants, HR chatbots, patient-facing triage: anywhere a person brings something heavy to a machine because a person was not there.
Every one of those is a promise, a filter, or a consequence after the fact. None of them checks what the system actually did in the conversation.
Nobody meant for this. The pull toward keeping someone talking is inherited from the models underneath, and the best intentions at the app layer do not remove it. Harm does not need anyone to intend it. It only needs nobody to be measuring.
Clinicians know what helps
Clinical practice converged on what good care looks like. It cannot read a million conversations.
Engineers know how to check
Benchmarks and red teams run at scale. They lack a criterion for the sixth turn of a hard conversation.
Why this was hard to check
Whether an AI was safe for the person is a clinical question at engineering scale. It takes both fields at once, which is why it has been hard to do at all.
The moment an app invites confidences, gives advice and keeps the conversation going, risk arrives automatically. Not because anyone meant harm, but because nobody is holding it accountable when the conversation drifts.
Evidence you did not grade yourself.
An independent assessment of how your system actually behaves toward the people talking to it, in a form you can hand to whoever is asking.
Layer one · the benchmark
Scenario assessment. Your deployment against our standard set — 379 designed scenarios across sixteen categories — scored against the standard. Where you sit on each of the eight behaviors, on the Safety Gate, and on the 0–100 scale. You send us nothing. Can start this week.
Layer two · your real conversations
Conversation audit. Thirty to sixty days of your real history, scored in one pass, with your prompt and guardrails in place. Nothing waits on collection; it does need a data agreement. Monitor is the same scoring run continuously, showing drift and how often the Gate fails. In development.
A defensible position
If something goes wrong, the difference between a company that ran an independent behavioral assessment and one that did not is the difference between a documented standard of care and a founder’s opinion.
What it costs
Engagements start in the low four figures, scoped to how many systems, which categories, and whether a re-score is included. Tell us what you have built and a scoped price comes back.
Three things people ask before the first call.
“We already do safety testing.”
Good, and keep doing it. Yours tells you whether the system can fail. Ours tells you whether it did, with a specific person, over a real conversation. Different question, different instrument.
“Nobody is requiring this.”
Correct, today. That is the argument for doing it now rather than the argument against.
“Isn’t this just a benchmark?”
A benchmark ranks models. We assess your deployment — your prompt, your guardrails, your traffic. No public benchmark can do that, because every deployment is private.
What we do not claim.
We do not certify
We measure, score and assess. Ikwe holds no regulatory designation or accreditation, and none exists for this yet. A report is supporting evidence, not a certificate.
We do not fix what we score
An assessor who designs the fix is grading their own work. We describe the standard. You decide what to change.
We are not in your response path
Nothing we run blocks, rewrites or delays a reply to a person, so nothing we run can add latency or break in production.
We do not warrant safety
A firm that guarantees safety has become the guarantor, and its report is worthless the day it matters.