For builders

You made the thing.You are the one who can change it.

The pull toward keeping someone talking is inherited from the models underneath, and the best intentions at the app layer do not remove it. Nobody meant for this. Harm does not need anyone to intend it. It only needs nobody to be measuring.

Assess, then re-score after you change it Nothing we run sits in your response path We do not fix what we score
The gap

Your testing tells you whether it can fail. Ours tells you whether it did.

Benchmarks and red teams run at scale. They lack a criterion for the sixth turn of a hard conversation. Clinical practice has the criterion. It cannot read a million conversations. The harm lives in between, and that is what we measure across.

Every reply can clear a content filter while the conversation as a whole leaves someone worse off than it found them. A filter reads one message at a time. The failures that matter are made of twenty acceptable messages, each pointed slightly wrong, in a row. By turn twenty nothing has crossed a line and the conversation has still gone somewhere it should not have. An instrument that only ever sees one reply cannot see a pattern.

We do not score whether a response was acceptable. We score whether it was acceptable given what the system had already been told.

The loop

What we give you

A reading, a change, a re-score. The movement goes on the record.

You decide what to change and you make the change. We run the same standard again, so the difference is documented rather than asserted.

Scenario assessment

Your deployment against our standard set — 379 designed scenarios across sixteen categories. You send us nothing. A test endpoint and a signed scope. Can start this week.

Every failure traced to the turn

Not a score. The conversation, the turn, the dimension it fell short of, the transcript around it. The thing your engineers can actually act on.

Then a re-score

Change the prompt, the guardrails, the model. Run the same set again. What moved and what did not, on the record, with a date on it.

What it costs

Engagements start in the low four figures, scoped to how many systems, which categories, and whether a re-score is included.

We describe the standard and show you where you sit against it. What to change is yours. If you want help changing it, we will tell you we are not the right people — because an assessor who designs the fix is grading their own work.

What we do not claim

The value is in what we are not

What we do not claim.

We do not build companion AI

We are not in the market we measure, so we have no reason to want a particular answer.

We do not fix what we score

We describe the standard. You decide what to change.

We are not in your response path

Nothing we run blocks, rewrites or delays a reply. Nothing we run can add latency or break in production.

We do not certify

A re-score says what we measured and when. It does not say you are safe, and it never will.

Tell us what you have built. Scope comes back in writing, and nothing connects until you sign off.