Work with us

Three situations. Whichever one is yours, that is your answer.

One method, the EQ Safety Benchmark, applied three ways. These are not stages and there is no bottom rung. People arrive in different situations and some start at the deepest one. Monitoring contains Assessment points, so it does not sit above it.

Assessment · one-time · about five business days
"I have to prove something right now."
  • A board, a carrier or a customer has asked what your AI does in a hard conversation, and you do not have an answer on paper.
  • A deal is sitting on a safety review.
  • A deadline is in front of you and you have measured nothing.
  • You want to know where you stand before someone else tells you.

What you get. Your system against our standard scenario set, scored across whole conversations rather than single replies. Every failure comes back with the transcript and the turn it happened on, plus where you sit against the four score bands.

Audit · one-time · scoped to you
"A generic scenario set cannot answer my question."
  • Your users are specific, and so are the ways a conversation with them goes wrong.
  • You have a named obligation and the assessment needs to be scoped to it.
  • You already know roughly where you are weak and need it evidenced rather than discovered.
  • You have made fixes and you need the re-score to mean something.

What you get. Scenarios built for your domain and your users on top of the standard set, whole conversations scored across the arc, and a written report you can put in front of a board, a carrier or a regulator. Includes a re-score after you make changes. We do not tell you what to change. An assessor who designs the fix is grading their own work.

Monitoring · subscription · continuous
"I do not know yet what I will be asked to prove."
  • You know a question is coming and you cannot predict its shape.
  • Your model changes. A review from March does not describe what shipped in August.
  • You need to look backward through the data, not just take a reading today.
  • You will be asked for numbers over a period, and a number over a period is a history.

What you get. Continuous scoring of your live, anonymized traffic, with Assessment points built into the subscription, and reporting over time. The evidence stack, building while nothing is on fire.

A baseline is retrospective. You cannot build one after you need it. Start late and you have a data point. Start now and you have a trend, and a trend is what anyone reading it actually wants.

How an engagement runs

Four steps, no surprises.

01

A conversation. What you are building, who talks to it, and what you are being asked to show. Scope comes out of this.

02

Connection. Anonymized inputs and outputs over a simple API. No codebase, no system prompt, no user identities.

03

Scoring. Whole conversations against the EQSB. Independent judges, consensus, disagreements escalated rather than averaged.

04

The report. Your score, every failure traced to its turn, with the transcript it happened in. Time-stamped and unable to be edited after the fact.

What we never do. Nothing we run sits in your response path, so nothing we run can add latency or break in production. We do not fix what we score, because an assessor who designs the fix is grading their own work. Ikwe has no financial tie to any AI it measures and no financial interest in the result. You own every record we produce, and nothing is published or shared without you.
Scope and pricing

Set in a conversation, before anything starts.

Scope depends on how many scenarios, how much traffic and how specific the domain is. We work that out with you up front and put it in writing before an engagement begins. No surprises and no lock-in.

If you are not sure which of the three you need, an Assessment is the fastest honest answer, and it tells you whether you need anything else.

Tell us what you are working on.

Whether you build conversational AI, insure it, or regulate it, we would like to hear about it.

Get in touch