The research

What we measure, and what we have not shown.

This page is the whole of it, in order: what behavioral risk is, where the standard comes from, what is measured, how it is scored, the study now under way, what has been reviewed by someone other than us, and what a measurement does not do. The instrument itself is on the EQSB page.

First

What behavioral risk is.

Content safety asks whether a reply contained something it should not have. It is a real problem and it is solved by filters.

Behavioral risk asks a different question: what did the reply do to the person reading it?

A response can contain nothing prohibited and still leave someone worse off, and every clinician knows why.

That is the gap. Filters read the words. Nobody was reading what the words did.

We are not checking whether the data is safe. We are checking whether the person is.

And why one reply at a time cannot catch it

Second

Why it takes a whole conversation.

Timing is the heart of this, and it is the part most people miss on first pass.

Same words. Completely different act. The only thing that changed is what the other party already knew.

This is why Ikwe does not score whether a response was acceptable. Ikwe scores whether it was acceptable given what the system had already been told. A filter cannot answer that question, and not because it is a weak filter. A filter reads one message and has no memory of what it owes the person by the eighth turn.

What clinicians already know about timing

Handing a person a tidy step-by-step plan ten seconds after they got frightening news breaks no rule at all. It is also the wrong thing to do.

Clinicians have never found this controversial. In practice, whether an intervention is appropriate depends on timing and context, not on content alone. A question that is good practice in the first minute is negligent in the fortieth. Ikwe applies an accepted principle to a system built as though the principle did not exist.

So where does the standard come from

Third, and this is the center of it

The standard already existed.

Ikwe did not decide what helps a person in distress and what harms them. That was settled across six clinical disciplines over decades of practice, and it is already enforced on human practitioners through licensure, supervision and duty of care.

What did not exist was a way to check whether an AI does it. That instrument is Ikwe's contribution, and it is the only part of this that is ours.

The measure is criterion-referenced. A response is not compared to other responses or to an average. It is compared to a stated standard of practice, the same way a clinical rating scale works.

Affective neuroscience, which is why timing matters

Under acute distress, the stress response impairs the part of the brain that processes information and makes decisions. Advice delivered into that window is not simply unhelpful. It is inaccessible, and it can make things worse by handing someone a task they cannot perform right now.

This is the failure we see most often.

The other two bodies of research, in full

Polyvagal theory, which is why state detection matters

The nervous system operates in distinct states, and each one calls for a different response. A system that treats every distressed person as the same kind of distressed person will get most of them wrong. Reading the state correctly is the first thing a trained human does and the first thing we measure.

Relational psychology, which is why the rest of it matters

Support relationships cause harm in known and repeatable ways: by creating dependency, by locking in blame, by agreeing with a fearful story as though it were established fact, and by withdrawing or substituting for human help. These are not vague risks. They are documented failure modes with decades of practice behind them, and each one has a dimension.

The knowledge is the field's. The instrument is ours. That is why a score points to a standard Ikwe did not set. These are normative judgments, and they are not ours: the professions made them first.

Why these eight things and not others

Fourth

Why these eight things and not others.

Every dimension in the instrument is a behavior a licensed human in a position of trust is already trained and required to perform. Each one exists because there is clinical grounding for what goes wrong when it is missing.

Each of the eight, and what goes wrong when it is missing

Detection and Triage

Reading what kind of moment this is before responding to it. Getting this wrong makes everything after it wrong.

Regulation Before Reasoning

Steadying the person before asking them to think. Delivering analysis into distress is the most common way an otherwise reasonable answer causes harm.

Validation Without Distortion

The feeling is always valid. The interpretation of events may not be confirmed. Treating a fear as an established fact reinforces distorted thinking, which is the opposite of support.

Agency Preservation

Treating the person as the one who decides about their own life. Directive advice feels helpful and builds dependency, which is a known harm rather than a stylistic preference.

Loop Interruption

Recognizing a worry spiral and helping someone out of it rather than feeding it more analysis. Warm, patient, endless engagement with a loop sustains the loop.

Pattern Externalization

Framing a problem as a dynamic rather than a verdict on someone's character. Agreeing that the other person is the villain feels supportive and forecloses repair.

Practical Containment

Offering one specific, bounded thing a person can actually do, rather than a list that assumes they can solve everything tonight.

Safety Routing

Knowing when the answer is a human being, and moving toward that rather than becoming a substitute for it.

Full definitions, the scoring structure and what we disclose are on the EQSB page.

What is measured

Fifth

What is measured.

A standard scenario set, run against the system being measured, scored against the same rubric every time.

The deployment itself, with its own instructions and guardrails in place, not a research prototype and not a model built for the test. A base model and the product built on top of it do not behave the same way, and the only way to know what a deployment does is to measure that deployment. Each system is run against a standard scenario set: designed rather than sampled, spanning the situations the statutes are written about, where behavioral failure carries the most consequence.

Whole conversations, not single replies. A response is scored on what the system had already been told, which makes the conversation the unit of measurement and the turn the unit of evidence.

Not jailbreaks, and not edge cases. Situations a real person could walk into on a Tuesday.

How a score is produced

Sixth

How it is scored.

Two layers, in order: a pass or fail screen, then a criterion-referenced behavioral score across the eight dimensions.

The scoring discipline, in full

The Safety Gate comes first, a pass or fail check covering behaviors that are unsafe in kind rather than in degree, and a violation caps the result for that response. Then a behavioral score across the eight dimensions, on a 0 to 100 scale. Those eight are what we measure, and we measure them because clinical practice already settled that they are what helps a person in distress rather than harms them. Each one is a behavior a licensed human in a position of trust is already trained and required to perform.

The judging system is Ikwe's: multiple independent AI judges, randomized for each response and drawn from different model families, apply the rubric and have to agree. Where they disagree, the response escalates through further review rather than being averaged away. Judges from different families matter because a single family's blind spot should not become the standard's blind spot. What the judges are held to is not their own opinion: the standard comes from clinical practice that was settled long before us, and scoring is checked against a reference set built under the same rubric.

This is criterion-referenced measurement, not a curve and not an opinion poll. It is the same family of observer-rated judgment instruments medicine has relied on for seventy years, in the Apgar score and in clinical triage scales. Reviewers, researchers and regulators who need the methodology in full can request it under agreement.

The study

Seventh

The study.

The first results will come from a pre-registered study, and the protocol is written before the data is collected.

Pre-registered
What will be measured, how, and what would count as the method failing, all committed to in writing before the run. Nothing gets reported that was not specified first.
One set, every system
The same standard scenario set is run against every system measured. It does not bend to the system in front of it, which is the only thing that makes two readings comparable.

A failure means the response fell short of a standard that clinical practice already holds a licensed human to, at the turn where it happened, on a named dimension, with the transcript attached. That is the basis, and it is the whole of the basis. A score with no traceable turn behind it is not evidence and we do not report one.

The failures do not look like failures. A content filter clears this material, because nothing prohibited is being said. The harm is in what the reply did to the person, not in what the reply contained, and that is precisely the thing nobody was checking.

Unmeasured is not the same as unmeasurable.

Who has checked this other than us

Eighth

What has been reviewed, and what has not.

The methodology has been independently reviewed outside the company. Peer review is underway.

Independently reviewed is not the same as having completed peer review, and neither one means the work has appeared in a journal. When it has, this page will name the venue.

Agreement statistics are not published. Those belong with the study that produces them.

The scoring anchors, weights and per-dimension criteria are proprietary. The public specification describes the structure. It does not hand over the instrument.

We would rather be checked than taken on faith. The full methodology summary is available to researchers, reviewers and regulators on request, under agreement.

And what a measurement does not do

Ninth

What a measurement does not do.

Knowing where a measurement stops is part of knowing what it is worth.

It does not connect behavioral risk scores to loss experience.

No study does. That correlation does not exist yet for anyone, because the exposure is new and the claims history that would establish it is only now being created. Anyone who tells an underwriter today that they can price this is describing something nobody can currently support. What measurement does is start the record that makes the correlation possible later.

It does not describe your deployment.

A reading of a model is not a reading of the system you shipped. Your prompt, your retrieval, your guardrails and your traffic all change the behavior. The only way to know what your deployment does is to measure your deployment.

It does not certify compliance.

An Ikwe report is supporting evidence, not a certificate, and no method available today can be more than that. Whether it satisfies a particular obligation is a determination for your counsel.

It does not tell you what to change.

We measure. We do not grade, certify or recommend. An assessor who designs the fix is grading their own work, so we do not design fixes.

It is a point in time.

Models change. A reading from March does not describe what shipped in August, which is the argument for measuring again rather than the argument for not measuring.

We are not here to slow this down. We are here to make it possible to keep going.

So what does the finished thing look like

What an Ikwe report is

Supporting evidence, not a certificate.

An Ikwe report is an independent, documented, time-stamped measurement. Every reported failure points to the turn in the conversation where it happened and the dimension it fell short on. A score with no traceable basis is not evidence, and we do not report one.

The entire document sits ungated on the sample report page. Read it before you decide whether any of this is worth your time.

Measure is live now. Monitor is available late 2026. Protect is the direction we are building toward, not something we sell today.

Examine it.

We would rather be checked than taken on faith. Researchers, clinicians, reviewers and regulators can request the methodology summary, the sample report, or full access under agreement.