The whole conversation, against eight behaviors clinical practice already established.
We did not invent the ethics. Six clinical disciplines did, over decades of practice. What we built is a way to check whether a machine actually does those things, turn after turn, and to show where it stopped.
The EQ Safety Benchmark scores a whole conversation rather than one reply at a time. A Safety Gate screens every response for ten coded failures that cap the score regardless of what else the response does well. Eight behavioral dimensions are then scored and combined into one result from 0 to 100. Systems run against a fixed set of 379 designed scenarios across sixteen categories; every system runs the same set.
Scoring uses several AI judges from different model families, assigned at random per response, each applying the same rubric. Disagreement escalates for review rather than being averaged away. Drawing judges from different families reduces the chance that one family’s blind spot becomes the instrument’s. It does not eliminate it.
What has been reviewed, and what has not.
Consistency, checked
Scoring is checked against a reference set scored under the same rubric. That shows the rubric is applied the same way twice.
Validity, not yet
A reference set built from an instrument cannot test that instrument. Testing it against independent human expert judgment is a separate step, and it is the study now under way.
The form is medicine’s
Defined behaviors, fixed anchors, trained raters — the form of the Apgar score and clinical triage scales. Those instruments earned their standing by publishing reliability and validating against outcomes. The EQ Safety Benchmark has not done that yet.
The bands are provisional
Boundaries were set by the rubric’s authors and have not been through a formal standard-setting procedure. A score of 85 means 85 on this scale, not a verified level of real-world safety.
What is published, what is shared, and what is protected.
Published
The dimension names and definitions, the Safety Gate categories, the scenario category list, the scoring bands, and the study pre-registration when it is filed.
Shared under agreement
Scenario text, scoring anchors, judge configuration and per-dimension criteria, for regulators, clinicians and researchers who request methodology access.
Protected
Dimension weights and the adjudication logic. An instrument whose weights are public can be optimized against, which is the failure it exists to detect.