The EQ Safety Benchmark

Independent behavioralrisk measurement for AI.

We score whole conversations against a human standard of care — how someone in a position of trust is expected to handle a person — and hand back an independent record. Not what your system can do under test. What it did, with a real person, turn by turn.

Two products, both deliverable today Nothing we run sits in your response path We do not fix what we score

Layer one · the benchmark

Scenario assessment. We run your system through our standard set — 379 designed scenarios across sixteen categories — and score every response against the standard. What comes back is where you sit: on each of the eight behaviors, on the Safety Gate, and on the scale from 0 to 100. You send us nothing. No data agreement, no privacy review, a test endpoint and a signed scope. This is the one that can start this week.

Layer two · your real conversations

Conversation audit. Thirty to sixty days of your actual history, scored in one pass, with your prompt and your guardrails in place. You already have that history, so nothing waits on collection — it does need a data agreement and a privacy review. Monitor is the same scoring run continuously against live traffic, showing where drift starts and how often the Safety Gate fails. In development.

What comes back, either way

A scored record naming the conversation, the turn, and the behavior that fell short, with the transcript around it. Yours to hand to a procurement officer, an underwriter, a board, or your counsel. Then a re-score after you make changes, so the movement is documented rather than asserted.

What it costs

Engagements start in the low four figures, scoped to how many systems, which categories, and whether a re-score is included. Tell us what you have built and a scoped price comes back. You do not have to sit through a discovery call to get a number.

You already have the data. We score your last sixty days rather than asking you to wait sixty days to collect it.

Ikwe holds no regulatory designation, approval or accreditation, and none exists for this yet. An Ikwe report is supporting evidence. It is not a certificate.

What it does for you

What it does for you

Before something happens, and after.

That is the whole point of having a measurement in place. Diagnostics while it is happening, so you can do something about it. And a record afterward, so you can see what you were doing.

Before · prevent it, and improve ahead of time

Run your system against the standard before anyone gets hurt. See where it falls short, on which turn, against which behavior. Change the prompt, the guardrails, the model. Run the same set again and the movement goes on the record. Measure does this today. Monitor — the same standard run continuously against live traffic, so drift is caught while nothing is on fire — is in development.

After · find out what happened, and fix it for next time

Something went wrong, or someone is asking whether it did. Score the real conversation history — thirty to sixty days that already exist — and get every failure traced to the turn it happened on, with the transcript around it. What went wrong, why, and where to change it. Time-stamped, from an outside measurer, and it cannot be edited after the fact.

Same instrument, same standard, both sides. The difference is whether you are reading the record before the question is asked or producing it after.

Why we do this

Why we do this

AI that helps people is worth keeping. Measurement is how we keep it.

Getting help today can mean months on a waiting list while things get worse. AI can be there in that gap, at any hour, at scale, for people the system has not reached. We are not here to slow that down.

Everything else a person turns to in distress is held to a standard. The hotline has protocols. The clinic has a licensing board. The counselor has supervision and a duty of care. None of that exists because anyone assumed they would cause harm. It exists because when someone is at their most vulnerable, good intentions have never been accepted as sufficient evidence.

This is where people are going now. Holding it to the same bar is the condition for keeping it. Unmeasured is not the same as unmeasurable.

Who this is for

Who this is for

The same instrument, from four sides of the table.

Each of these is a different reason to need an independent record of how a system actually behaved. Pick the one that is yours.

You built it →

You make the model or the product. The pull toward keeping someone talking is inherited from the models underneath, and you are the one who can change what the system does. An assessment, then a re-score after you do.

You run it →

You deployed a conversational system in your product — HR, benefits, intake, tutoring, care. You may not have built it. You are the one who answers for it. Evidence you did not grade yourself.

You insure it →

E&S and specialty carriers writing an exposure with severity but no frequency, correlated across a book that bought the same three underlying models. A pre-loss signal where none exists.

You are answering for it →

A complaint, a civil investigative demand, a regulator’s inquiry, or diligence on a system you bought rather than built. A time-stamped record that survives challenge.

You are checking it →

Regulators, researchers and clinicians who want to examine the method rather than take our word for it. What is published, what is shared under agreement, what the study will show.

You are just looking →

Not in any of those seats yet, or all of them at once. Start with the demo: pick a situation, walk the conversation, watch where the score moves. Then read the instrument, the research, or a sample report at your own pace.

Why this is arriving now

Why now

The asks have started.

Enterprise buyers want assurance in procurement. Insurers are beginning to ask what the exposure looks like. Boards ask what happens if.

A dozen-odd states enacted conversational-AI statutes in 2026, and the list is still growing. Every one of them requires protocols for responding to expressions of self-harm — which somebody eventually has to demonstrate rather than assert. California tells operators to use “evidence-based methods for measuring suicidal ideation” and never defines the term. Every operator in scope is choosing a method right now, by default, and will have to defend that choice later.

Nobody is requiring an independent behavioral assessment today. That is the reason to do it now rather than the reason not to. The companies that documented food safety before it was mandatory were not the ones who suffered when it became mandatory.

The full map: which states, which dates, and what each statute actually asks for.

Get in touch

Start here

Tell us what you are working on.

Whether you build conversational AI, insure it, or regulate it, Ikwe would like to hear about it. A reply comes back within two business days.

Prefer email? Write to hello@ikwe.ai. Researchers, clinicians and regulators can request methodology access at ikwe.ai/access.

What happens next. Every note goes to a person, not a queue, and a reply comes back within two business days. If a Measure engagement is the next step, scope is agreed in writing, and nothing connects until you sign off.