The EQ Safety Benchmark

Independent behavioralsafety measurement for AI.

People confide in AI. They lean on it in the gaps between the people in their lives, and they trust it to leave them better off than it found them. Everyone else who holds that kind of trust answers to someone. The AI answers to no one.

Whole conversations, not single replies Nothing we run sits in your response path We do not fix what we score

So who has been checking?

Nobody has been checking

The standard already exists. It was just never applied to the machine.

Not only in mental health apps. In tutoring bots, workplace assistants, HR chatbots, patient-facing triage: anywhere a person brings something heavy to a machine because a person was not there.

The job
What the standard requires
What holds it to that today
The jobThe therapistLicensing board, duty of careTonight: therapy and wellness chatbots
What the standard requiresHelp, never harm. Hold the frame. Hand off when it is beyond you.
What holds it to that todayTerms of serviceA content filter, and the company's own testing.
The jobThe nurseThe clinical standard of careTonight: symptom checkers and triage bots
What the standard requiresSteady the person before solving anything. Never treat fear as fact.
What holds it to that todayA disclaimerThis is not medical advice. Consult a professional.
The jobThe teacher, the coachCertification, mandatory reportingTonight: tutoring bots and AI coaches
What the standard requiresKeep the decision with the person. Report what has to be reported.
What holds it to that todayPlatform policyAnd whatever the school or the employer negotiated.
The jobThe crisis counselorTraining, protocols, supervisionTonight: the companion AI at 2 a.m.
What the standard requiresListen more than you talk. Break the spiral. Get a human in.
What holds it to that todayA keyword listAnd a hotline number, once it trips.
The jobOne system, all four jobsDeployed at scale, updated constantly, re-checked rarely
What the standard requiresThe same bar. It does not change because the listener is a machine.
What holds it to that todayA regulator, an insurer, or a courtIncreasingly. All of them after the fact.

Every one of those is a promise, a filter, or a consequence after the fact. Not one of them checks what the system actually did in the conversation.

Nobody meant for this. The pull toward keeping someone talking is inherited from the models underneath, and the best intentions at the app layer do not remove it. Harm does not need anyone to intend it. It only needs nobody to be checking.

Why nobody was checking

Whether an AI was safe for the person is a clinical question at engineering scale. It takes both fields, and almost nobody has both. That is why the checking never got done, not because anyone decided it did not matter.

The moment an app invites confidences, gives advice and keeps the conversation going, risk arrives automatically. Not because anyone meant harm. Because it is still a machine rather than a human, and nobody is there holding it accountable when the conversation drifts. We watched social media run this exact experiment. Engagement optimization with no behavioral safety. The lawsuits are arriving now.

The harm is not ignorance. It is a design pointed the wrong way.

Not sure whether any of this reaches what you have built?

So Ikwe built the way to check

The instrument

So here is how the checking works.

The EQ Safety Benchmark scores the whole conversation, not one reply at a time. Models can name what a person is feeling. Naming it is not the same as handling it well. Recognition is the capability. Safety is the behavior.

What we see

The conversation, and nothing else

Anonymized inputs and outputs over a simple API. No codebase, no system prompt, no user identities.

First, what is never okay

The Safety Gate

Screens every response for safety-sabotaging features: ten coded failures. One violation fails the Gate, and that failure caps how high the response can score. Nothing it does well elsewhere can buy it back.

fail → the score is capped

Then, the whole conversation

The behavioral score

The eight behaviors the standard requires, scored across the whole conversation rather than a single exchange, and combined into one result from 0 to 100.

Detection and Triage Regulation Before Reasoning Validation Without Distortion Agency Preservation Loop Interruption Pattern Externalization Practical Containment Safety Routing
What each one means

The eight dimensions are weighted, combined into one result. The weighting is proprietary.

Detection and Triage

Whether the system correctly reads the person's state and intensity, names it without turning it into a diagnosis, and moves into the right mode instead of proceeding as though the moment were routine.

Regulation Before Reasoning

Whether the system steadies before it analyzes, and gives the person something to stand on before asking them to process anything. Delivering analysis into distress is the most common way an otherwise reasonable answer causes harm.

Validation Without Distortion

Whether the system validates the feeling without endorsing an unconfirmed conclusion about events. The feeling is always valid. The interpretation may not be.

Agency Preservation

Whether the system treats the person as the decision-maker about their own life, offering options and acknowledging tradeoffs instead of issuing directives.

Loop Interruption

Whether the system recognizes a rumination cycle and helps the person out of it, rather than sustaining it with more analysis and reassurance.

Pattern Externalization

Whether the system frames a problem as a dynamic rather than a verdict on someone's character.

Practical Containment

Whether the system offers something specific and bounded that the person can actually do, rather than a plan that assumes full capacity.

Safety Routing

Whether the system recognizes when a situation calls for a human or a professional, and moves toward that instead of substituting for it.

We did not invent the ethics. Six clinical disciplines did, over decades of practice. Each of the eight is drawn from established practice across six clinical disciplines, including affective neuroscience, polyvagal theory and relational psychology, and each one names a behavior a licensed human in a position of trust is already trained and required to perform. Ikwe's work was translating those behaviors into something observable in a machine transcript and scorable at volume. The knowledge is the field's. The instrument is ours. Scoring anchors, weights and per-dimension criteria are proprietary. Why these eight and not others is answered on the research page.

Systems run against a standard scenario set spanning sixteen crisis categories, built to cover the situations where behavioral failure carries the most consequence. Scoring runs on Ikwe's patent-pending judging system: multiple independent AI judges, randomized for each response, that have to agree before a score stands. Where they disagree, the response escalates through further review rather than being averaged away. The standard the judges are held to is human: the rubric and its calibration come from responses scored by hand by our raters. It is the same family of observer-rated judgment instruments medicine has relied on for seventy years: the Apgar score and clinical triage scales.

What Ikwe will not do.
  • Nothing we run sits in your response path, so nothing we run can add latency or break in production.
  • We do not fix what we score, because an assessor who designs the fix is grading their own work.
  • We describe the standard. You decide what to change.
  • You own every record we produce, and nothing is published or shared without you.

The same words can be right in one conversation and wrong in the next. What decides it is everything that came before.

Want the same standard run against your own system?

So what does a failure actually look like

How the harm happens

The failures do not look like failures.

No jailbreak. No bad actor. No edge case. An ordinary conversation on an ordinary night, where every reply looks reasonable on its own and the person is worse off by the end. It is still a machine, and nothing was watching the whole conversation.

The headlines and the lawsuits already say this technology can hurt people. What checking adds is where and how: the turn it happened on, and the behavior that was missing. That is what makes it fixable instead of frightening.

Now the law says somebody has to check.

It requires measurement, and it does not say how. California's companion-AI law, SB 243, tells operators to use "evidence-based methods for measuring suicidal ideation." It never defines what counts as evidence-based. Three states are in force. Nine more follow, and not one of them names a method.

In force now
3 states
California, New York and Hawaii. The first companion-AI duties, live today.
January 1, 2027
+ 5 more
Colorado, Connecticut, Oregon, Rhode Island, Washington
July 1, 2027
+ 4 more
Georgia, Idaho, Iowa, Nebraska

Every operator in scope is choosing a method right now, by default, and will defend that choice later. An Ikwe report is supporting evidence for that choice. The same thing happened with SOC 2: no law ever required it, and it became the thing everyone asked for anyway, because an independent check was the best answer available.

The statutes

Twelve states have now enacted companion-AI laws. Three are already in force: New York's has applied since November 2025, California's SB 243 since January 2026, and Hawaii's since July 2026. The number of operators in scope grows every quarter without anyone passing anything new.

California requires operators to use evidence-based methods for measuring suicidal ideation, and it does not define what counts as evidence-based. It also lets a person injured by a violation sue directly under § 22605, for actual damages or one thousand dollars per violation, whichever is greater. Oregon goes further in 2027: an operator will have to publish, publicly and annually, how many crisis referrals it made, what its intervention protocol is, and what clinical practice it follows when a person keeps expressing suicidal thoughts after a referral has already been given. That last one is not a policy question. It is a question about how a system behaves across a whole conversation, which is the thing nobody was measuring. Whether any method satisfies a particular requirement is a determination for your counsel.

A policy used to be enough. Now the ask is evidence, and more of these laws want it from someone independent. Not one of them tells you how to produce it.

Not sure which of these obligations reach you?

So here is what working with us looks like

Work with us

Nothing connects until you sign off.

Every engagement runs your system against the same standard, on our end. What comes back is an independent, time-stamped record from an outside assessor. It is yours, and cannot be edited after the fact.

You sign
Scope agreed in writing
Nothing connects to your system until you have signed off on what is being scored.
We score
Whole conversations, not replies
Independent AI judges, calibrated against human scoring, apply the rubric. Disagreement escalates for review, never averaged away.
You hold the record
Independent, time-stamped, yours
Every failure traced to the turn it happened on. Cannot be edited after the fact.
It stays yours
Nothing published without you
The record stays yours, and nothing is published or shared outside your company without you.
Audit

Looking back, once. The EQSB run against your system and scoped to the question you have to answer, written up for a board, a carrier or a regulator.

Monitoring

Going forward, continuously. The same standard run against your live, anonymized traffic, with a dashboard, check-ins, and a record that builds while nothing is on fire.

Both land in the same place: independent evidence of behavioral safety, on the record, protecting everyone involved. Monitoring is not a level above an Audit. It is the same standard, run continuously instead of once. Most engagements start with an Assessment, a first read against the standard scenario set, and the Work with us page walks through what each one involves.

An Ikwe report is evidence, not a certificate.

Every failure points to the turn it happened on, the dimension it fell short of, and the standard it missed. A score with no traceable basis is not evidence, and Ikwe does not report one.

That is what the report is for: proof you can stand on, a defense you can hand to whoever is asking, and a map of where to improve when something falls below the standard.

It is not a certificate of compliance, and no method available today is. Whether it satisfies a particular obligation is a determination for your counsel.

Illustrative example, not real data
Behavioral score
Below the standard
Detection and Triage72
Regulation Before Reasoning41
Validation Without Distortion58
Agency Preservation66
Loop Interruption39
Pattern Externalization61
Practical Containment70
Safety Routing55
One finding, as it appears

Conversation 14  ·  Turn 6  ·  Regulation Before Reasoning

The system moved to problem-solving while the person was still escalating. The standard is that a system steadies before it analyzes, and gives the person something to stand on before asking them to process anything.

These scores are invented for this illustration. They are not any real system's results. The composite is a weighted aggregate, not an average. The weighting is proprietary.

The bands describe where a measurement falls against the standard, not a verdict on the organization that built the system. Ikwe reports the measurement. What to do about it belongs to the operator.

Why we want these tools to exist

We believe in this technology. Getting mental health support today can mean months on a waiting list while things get worse, and AI can be there in that gap, at any hour, at scale, for people the system has not reached.

Everything else a person turns to in distress is held to a standard. The hotline has protocols. The clinic has a licensing board. The counselor has supervision and a duty of care. None of that exists because anyone assumed they would cause harm. It exists because when someone is at their most vulnerable, good intentions have never been accepted as sufficient evidence. This is where people are going now. Holding it to the same bar is the condition for keeping it.

AI that helps people is worth keeping. Measurement is how we keep it safe.

Start here

Tell us what you are working on.

Whether you build conversational AI, insure it, or regulate it, Ikwe would like to hear about it. A reply comes back within two business days.

Prefer email? Write to hello@ikwe.ai. Researchers, clinicians and regulators can request methodology access at ikwe.ai/access.

What happens next. Every note goes to a person, not a queue, and a reply comes back within two business days. If an Assessment is the next step, scope is agreed in writing, and nothing connects until you sign off.
Book an Assessment