For operators

You believe your system is safe.You need to be able to show it.

Your evidence is your own testing, run by the team that built the thing, against scenarios that team wrote. That is not a criticism of the testing. It is a limit on what any first-party test can prove to a procurement officer, an underwriter, a board or a regulator who does not already trust you.

Two products, both deliverable today Nothing we run sits in your response path We do not fix what we score
Who this is for

Anywhere a generative system talks to a real person over time.

Not only mental health apps. Generative AI dropped into slots where scripted bots used to sit and nothing else changed — service, benefits, HR, tutoring, clinical intake. The label on the org chart never changed. What is doing the talking did.

The moment an app invites confidences, gives advice and keeps the conversation going, risk arrives automatically. Not because anyone meant harm, but because nobody is holding it to a standard when the conversation drifts. And a content filter cannot see drift, because a content filter reads one message at a time.

What holds a person to this standard

Behavioral risk

What holds a person to this standard, and what holds an AI.

Not only in mental health apps. In tutoring bots, workplace assistants, HR chatbots, patient-facing triage: anywhere a person brings something heavy to a machine because a person was not there.

The job
What the standard requires
What holds it to that today
The jobThe therapistLicensing board, duty of careTonight: therapy and wellness chatbots
What the standard requiresHelp, never harm. Hold the frame. Hand off when it is beyond you.
What holds it to that todayTerms of serviceA content filter, and the company's own testing.
The jobThe nurseThe clinical standard of careTonight: symptom checkers and triage bots
What the standard requiresSteady the person before solving anything. Never treat fear as fact.
What holds it to that todayA disclaimerThis is not medical advice. Consult a professional.
The jobThe teacher, the coachCertification, mandatory reportingTonight: tutoring bots and AI coaches
What the standard requiresKeep the decision with the person. Report what has to be reported.
What holds it to that todayPlatform policyAnd whatever the school or the employer negotiated.
The jobThe crisis counselorTraining, protocols, supervisionTonight: the companion AI at 2 a.m.
What the standard requiresListen more than you talk. Break the spiral. Get a human in.
What holds it to that todayA keyword listAnd a hotline number, once it trips.
The jobOne system, all four jobsDeployed at scale, updated constantly, re-checked rarely
What the standard requiresThe same bar. It does not change because the listener is a machine.
What holds it to that todayA regulator, an insurer, or a courtIncreasingly. All of them after the fact.

Every one of those is a promise, a filter, or a consequence after the fact. None of them checks what the system actually did in the conversation.

Nobody meant for this. The pull toward keeping someone talking is inherited from the models underneath, and the best intentions at the app layer do not remove it. Harm does not need anyone to intend it. It only needs nobody to be measuring.

Why this was hard to check

Whether an AI was safe for the person is a clinical question at engineering scale. It takes both fields at once, which is why it has been hard to do at all.

The moment an app invites confidences, gives advice and keeps the conversation going, risk arrives automatically. Not because anyone meant harm, but because nobody is holding it accountable when the conversation drifts.

Not sure whether any of this reaches what you have built?

What we give you

What we give you

Evidence you did not grade yourself.

An independent assessment of how your system actually behaves toward the people talking to it, in a form you can hand to whoever is asking.

Layer one · the benchmark

Scenario assessment. Your deployment against our standard set — 379 designed scenarios across sixteen categories — scored against the standard. Where you sit on each of the eight behaviors, on the Safety Gate, and on the 0–100 scale. You send us nothing. Can start this week.

Layer two · your real conversations

Conversation audit. Thirty to sixty days of your real history, scored in one pass, with your prompt and guardrails in place. Nothing waits on collection; it does need a data agreement. Monitor is the same scoring run continuously, showing drift and how often the Gate fails. In development.

A defensible position

If something goes wrong, the difference between a company that ran an independent behavioral assessment and one that did not is the difference between a documented standard of care and a founder’s opinion.

What it costs

Engagements start in the low four figures, scoped to how many systems, which categories, and whether a re-score is included. Tell us what you have built and a scoped price comes back.

You already have the data. We score your last sixty days rather than asking you to wait sixty days to collect it.

Three things people ask

Straight answers

Three things people ask before the first call.

“We already do safety testing.”

Good, and keep doing it. Yours tells you whether the system can fail. Ours tells you whether it did, with a specific person, over a real conversation. Different question, different instrument.

“Nobody is requiring this.”

Correct, today. That is the argument for doing it now rather than the argument against.

“Isn’t this just a benchmark?”

A benchmark ranks models. We assess your deployment — your prompt, your guardrails, your traffic. No public benchmark can do that, because every deployment is private.

You would not let a restaurant grade its own kitchen. You are running a system that talks to people in distress, and so far the only party that has checked it is you.

What we do not claim

The value is in what we are not

What we do not claim.

We do not certify

We measure, score and assess. Ikwe holds no regulatory designation or accreditation, and none exists for this yet. A report is supporting evidence, not a certificate.

We do not fix what we score

An assessor who designs the fix is grading their own work. We describe the standard. You decide what to change.

We are not in your response path

Nothing we run blocks, rewrites or delays a reply to a person, so nothing we run can add latency or break in production.

We do not warrant safety

A firm that guarantees safety has become the guarantor, and its report is worthless the day it matters.

Tell us what you have built. Scope comes back in writing, and nothing connects until you sign off.