Work with us

One way in. Two ways to run it.

One method, the EQ Safety Benchmark. An Audit runs it once, looking back. Monitoring runs it continuously, going forward. Every engagement opens with an Assessment: a first read of your system against the standard scenario set, so you know where you stand before deciding which of the two you need.

Every engagement opens with an Assessment. Your system against the standard scenario set, scored across whole conversations rather than single replies. Every failure comes back with the transcript and the turn it happened on, plus where you sit against the four score bands. It is how you find out where you stand, and it is where the record starts.
Audit · once · looking back
When you have something to answer for now
  • A board, a carrier, a customer or a regulator has asked what your AI does in a hard conversation, and you do not have an answer on paper.
  • A deal is sitting on a safety review, or a named obligation needs the reading scoped to it.
Also this, and what you get

What you get. The EQSB run against your system, scoped to the question you have to answer, whole conversations scored across the arc, and a written report you can put in front of whoever is asking. We do not tell you what to change. An assessor who designs the fix is grading their own work.

Monitoring · continuous · going forward
When you do not yet know what you will be asked to prove
  • You know a question is coming and you cannot predict its shape.
  • Your model changes. A reading from March does not describe what shipped in August.
Also this, and what you get

You will be asked for a picture over a period, and a picture over a period is a history.

What you get. The same standard run against your live, anonymized traffic, with a dashboard, check-ins on a regular cadence, a named person on your account, and reporting over time. The record builds while nothing is on fire.

Whether you are answering for something now or watching for what is coming, both land in the same place: independent evidence of behavioral safety, protecting everyone involved.

These are not stages and there is no bottom rung. Monitoring is not a level above an Audit. It is the same standard, run continuously instead of once, and people arrive needing either one.

Who this is for

The same instrument, from either side of the table.

If counsel is involved

For counsel.

If you are responding to a complaint, a civil investigative demand, or a regulator's inquiry, an engagement can be scoped with your counsel from the start, including what is produced, how work product is handled, and the timeline the matter requires. Start that conversation through the contact page and say counsel is involved.

If the AI is a vendor's

You bought it. You did not build it.

If the conversational AI in your product is a vendor's, an Assessment can be commissioned by you and run on the system you are buying, as part of vendor risk review. Same instrument, same report, commissioned by the buyer rather than the builder.

The whole engagement, in one row

How an engagement runs

No surprises.

Scope is agreed in writing before anything starts. Nothing connects to your system until you have signed off on what is being scored. The record that comes back is yours.

How scope is set

Scope depends on how many scenarios, how much traffic and how specific the question is. We work that out with you up front and put it in writing before an engagement begins. No surprises and no lock-in.

If you are not sure which of the two you need, the Assessment answers it. It tells you where you stand, and whether a one-time reading is enough or the question is going to keep coming back.

What Ikwe will not do. Nothing we run sits in your response path, so nothing we run can add latency or break in production. We do not fix what we score, because an assessor who designs the fix is grading their own work. We describe the standard. You decide what to change. You own every record we produce, and nothing is published or shared without you.

Two situations worth naming

The engagement, in scope

You approve the scope before anything runs.

You give us
A connection, not your code

For an Assessment or Audit, our scenario set runs against your system from our end, over a simple API. For Monitoring, your anonymized live traffic. Never your codebase, your system prompt, or your users' identities.

How it is scored
Whole conversations

Every response passes the Safety Gate screen, then the eight-dimension standard, scored by Ikwe's patent-pending judging system: independent AI judges, calibrated against human scoring, that have to agree. Disagreement escalates for review, never averaged away.

Every score shows its work
Traced to the turn

Where you sit against the four score bands, with every failure traced to the turn it happened on and the transcript around it.

What comes back
Scored reporting

An Assessment or Audit returns a scored report. Monitoring returns a live-scoring dashboard, check-ins on a regular cadence, and a named person on your account.

Start here

Whichever one is yours, it starts the same way.

Thirty minutes. We will tell you where to start, scope it in writing, and send pricing. Nothing connects until you sign off, and the record that comes back is yours. If none of this is right for you yet, we will say so.

Usually it is the head of product, the general counsel, or whoever owns risk who brings us in. Bring whoever will have to answer the question, and whoever signs.

01

You do not have an answer on paper

Someone has started asking what your AI does in a hard conversation. An Assessment tells you where you stand, and where the record starts.

02

Something specific has been asked

A board, a carrier or a regulator has a named question. An Audit is looking back once, scoped to it, written up for whoever is asking.

03

You cannot predict the question yet

Your model changes and a reading from March does not describe what shipped in August. Monitoring is the same standard, run continuously.

A model that changes needs a reading that changes with it. A record built over time shows a trend rather than a single point.