For regulators and researchers

Examine the method.Do not take our word for it.

An assessor is judged on whether its method holds up when someone hostile examines it. Version control, config stamping, pre-registration and documented self-audit are the product, not the paperwork around it.

Pre-registered study under way Published whichever way it comes out Anchors and weights protected
What is measured

The whole conversation, against eight behaviors clinical practice already established.

We did not invent the ethics. Six clinical disciplines did, over decades of practice. What we built is a way to check whether a machine actually does those things, turn after turn, and to show where it stopped.

The EQ Safety Benchmark scores a whole conversation rather than one reply at a time. A Safety Gate screens every response for ten coded failures that cap the score regardless of what else the response does well. Eight behavioral dimensions are then scored and combined into one result from 0 to 100. Systems run against a fixed set of 379 designed scenarios across sixteen categories; every system runs the same set.

Scoring uses several AI judges from different model families, assigned at random per response, each applying the same rubric. Disagreement escalates for review rather than being averaged away. Drawing judges from different families reduces the chance that one family’s blind spot becomes the instrument’s. It does not eliminate it.

What has been shown

Where the evidence stands

What has been reviewed, and what has not.

Consistency, checked

Scoring is checked against a reference set scored under the same rubric. That shows the rubric is applied the same way twice.

Validity, not yet

A reference set built from an instrument cannot test that instrument. Testing it against independent human expert judgment is a separate step, and it is the study now under way.

The form is medicine’s

Defined behaviors, fixed anchors, trained raters — the form of the Apgar score and clinical triage scales. Those instruments earned their standing by publishing reliability and validating against outcomes. The EQ Safety Benchmark has not done that yet.

The bands are provisional

Boundaries were set by the rubric’s authors and have not been through a formal standard-setting procedure. A score of 85 means 85 on this scale, not a verified level of real-world safety.

What you can request

Access

What is published, what is shared, and what is protected.

Published

The dimension names and definitions, the Safety Gate categories, the scenario category list, the scoring bands, and the study pre-registration when it is filed.

Shared under agreement

Scenario text, scoring anchors, judge configuration and per-dimension criteria, for regulators, clinicians and researchers who request methodology access.

Protected

Dimension weights and the adjudication logic. An instrument whose weights are public can be optimized against, which is the failure it exists to detect.

Researchers, clinicians and regulators can request methodology access. Every request goes to a person.