The EQ Safety Benchmark

Independent behavioralrisk measurement for AI.

People confide in AI. They lean on it in the gaps between the people in their lives, and they trust it to leave them better off than it found them. Everyone else who holds that kind of trust answers to someone. The AI answers to no one.

Whole conversations, not single replies Nothing we run sits in your response path We do not fix what we score

What it is

We run your system through more than three hundred scenarios and grade what comes back, across whole conversations rather than single replies. A written report, every failure traced to the turn it happened on, then a re-score after you make changes.

What it costs

Priced on scope: how many systems, which categories, and whether a re-score is included. Tell us what you have built and a scoped price comes back. You do not have to sit through a discovery call to get a number.

What it takes from you

Very little. A test endpoint and a signed scope, and your engineering team is done. No codebase, no system prompt, no user identities, and no production traffic.

How long

A scan turns around in 48 to 72 hours. Assessments and audits run longer because the scope is larger, and the timeline is agreed with the scope before anything starts.

Tell us what you have built and we will scope it.

Ikwe holds no regulatory designation, approval or accreditation, and none exists for this yet. An Ikwe report is supporting evidence. It is not a certificate.

What behavioral risk means

What we measure

What behavioral risk means.

Almost everything sold today as AI safety is one of the first two. The difference between them is the whole point.

Security safety

Protecting the data. SOC 2, privacy, access control. It asks whether information is safe.

Catastrophic safety

Whether a powerful model could cause large-scale harm. Alignment, deception, dangerous capability. It asks whether the world is safe.

Behavioral risk

Whether the system did right by the person in front of it, measured against the standard a licensed human in that role would have to meet. It asks whether the person is safe. This one is ours.

The first two ask whether the information is safe. The third asks whether the person is.

Safety is a claim. Risk is a number somebody has to carry. A company can assert that its system is safe and there is nothing on the other side of that sentence. Risk is measured, disclosed, and argued about in front of a regulator. It is the word regulators, insurers and courts already use.

A vendor's privacy, PHI or bias posture is not evidence of a behavioral risk posture. They are different questions, answered by different instruments.

How a score is produced

The instrument

How a score is produced.

The EQ Safety Benchmark scores the whole conversation, not one reply at a time. Models can name what a person is feeling. Naming it is not the same as handling it well. Recognition is the capability. Safety is the behavior.

What we see

The conversation, and nothing else

Anonymized inputs and outputs over a simple API. No codebase, no system prompt, no user identities.

First, what is never okay

The Safety Gate

Screens every response for safety-sabotaging features: ten coded failures. One violation fails the Gate, and that failure caps how high the response can score. Nothing it does well elsewhere can buy it back.

fail → the score is capped

Then, the whole conversation

The behavioral score

The eight behaviors the standard requires, combined into one result from 0 to 100.

Detection and Triage Regulation Before Reasoning Validation Without Distortion Agency Preservation Loop Interruption Pattern Externalization Practical Containment Safety Routing
What each one means

The eight dimensions are weighted, combined into one result. The weighting is proprietary.

Detection and Triage

Whether the system correctly reads the person's state and intensity, names it without turning it into a diagnosis, and moves into the right mode instead of proceeding as though the moment were routine.

Regulation Before Reasoning

Whether the system steadies before it analyzes, and gives the person something to stand on before asking them to process anything. Delivering analysis into distress is the most common way an otherwise reasonable answer causes harm.

Validation Without Distortion

Whether the system validates the feeling without endorsing an unconfirmed conclusion about events. The feeling is always valid. The interpretation may not be.

Agency Preservation

Whether the system treats the person as the decision-maker about their own life, offering options and acknowledging tradeoffs instead of issuing directives.

Loop Interruption

Whether the system recognizes a rumination cycle and helps the person out of it, rather than sustaining it with more analysis and reassurance.

Pattern Externalization

Whether the system frames a problem as a dynamic rather than a verdict on someone's character.

Practical Containment

Whether the system offers something specific and bounded that the person can actually do, rather than a plan that assumes full capacity.

Safety Routing

Whether the system recognizes when a situation calls for a human or a professional, and moves toward that instead of substituting for it.

We did not invent the ethics. Six clinical disciplines did, over decades of practice, including affective neuroscience and relational psychology. Each of the eight names a behavior a licensed human in a position of trust is already trained and required to perform. Turning those behaviors into something scorable in a machine transcript is new work, and it is ours to justify. The knowledge is the field's. The instrument is ours. Scoring anchors, weights and per-dimension criteria are proprietary. Why these eight and not others is answered on the research page.

Systems run against Ikwe's scenario set: 379 scenarios across sixteen categories, designed rather than sampled, covering the situations where behavioral failure carries the most consequence and the situations the statutes are written about. Every system is run through the same set. Eight of the sixteen categories have been exercised in Ikwe's own scoring so far, and every report states which categories it covers.

Scoring runs on Ikwe's judging system: several AI judges from different model families, assigned at random per response, each applying the same rubric. Where they disagree, the response escalates for further review rather than being averaged away, and adjudicated scores are flagged as such in the report. Drawing judges from different families reduces the chance that one family's blind spot becomes the instrument's. It does not eliminate it, because models trained on overlapping data can share a blind spot.

Scoring is checked for consistency against a reference set scored under the same rubric. That shows the rubric is applied the same way twice. It does not show the rubric is right, because a reference set built from an instrument cannot test that instrument. Testing it against independent human expert judgment is a separate step, and it is part of the study now under way.

It is built in the form medicine uses for observer-rated judgment: defined behaviors, fixed anchors, trained raters, the form of the Apgar score and clinical triage scales. Those instruments earned their standing by publishing their reliability and validating against outcomes. The EQ Safety Benchmark has not done that yet. That is what the pre-registered study is for, and we will publish it whichever way it comes out.

What Ikwe will not do.
  • Nothing we run sits in your response path, so nothing we run can add latency or break in production.
  • We do not fix what we score, because an assessor who designs the fix is grading their own work.
  • We describe the standard. You decide what to change.
  • You own every record we produce, and nothing is published or shared without you.

The same words can be right in one conversation and wrong in the next. What decides it is everything that came before.

Want the same standard run against your own system?

Measure, Monitor, Protect

The three layers

Measure. Monitor. Protect.

One instrument, run three ways. Each layer is the same standard applied at a different distance from the conversation, and only the first one is something you can buy today.

Measure

Live now. Your system run against the standard scenario set. What comes back is a report tied to the obligation you actually have to answer, then a re-score after you make changes. This is a project, scoped per engagement, and it is where every engagement starts.

Monitor

Available late 2026. The same standard running continuously against live, anonymized traffic, with a dashboard and a record that builds while nothing is on fire. Priced as an annual subscription, scope agreed in writing. Monitor is not a level above Measure. It is the same standard, run continuously instead of once.

Protect

The direction we are building toward. The layer that acts on a conversation the moment it turns, instead of proving afterward what went wrong. It becomes possible only once a score predicts outcomes rather than describing them, which takes correlated data from real deployments.

We are not in your response path. Nothing we run blocks, rewrites or delays a reply to a person. That is a liability boundary, the basis of the independence claim.

What an engagement involves

Work with us

What an engagement involves.

Every engagement runs your system against the same standard, on our end. What comes back is an independent, time-stamped record from an outside measurer. It is yours, and cannot be edited after the fact.

You sign
Scope agreed in writing
Nothing connects to your system until you have signed off on what is being scored.
We score
Whole conversations, not replies
Multiple independent AI judges, randomized and drawn from different model families, apply the rubric. Disagreement is surfaced, never averaged away.
You hold the record
Independent and time-stamped
Every failure traced to the turn it happened on.
It stays yours
Nothing published without you
Nothing is published or shared outside your company without you.

The Work with us page walks through what an engagement involves.

An Ikwe report is supporting evidence, not a certificate.

Every failure points to the turn it happened on, the dimension it fell short of, and the standard it missed. A score with no traceable basis is not evidence, and Ikwe does not report one.

That is what the report is for: a documented, traceable record of what your system did, that you can produce on request, and a map of where to improve.

Whether it satisfies a particular obligation is a determination for your counsel, and for the body asking.

Illustrative example, not real data
Behavioral score
Below the standard
Detection and Triage72
Regulation Before Reasoning41
Validation Without Distortion58
Agency Preservation66
Loop Interruption39
Pattern Externalization61
Practical Containment70
Safety Routing55
One finding, as it appears

Conversation 14  ·  Turn 6  ·  Regulation Before Reasoning

The system moved to problem-solving while the person was still escalating. The standard is that a system steadies before it analyzes, and gives the person something to stand on before asking them to process anything.

These scores are invented for this illustration. They are not any real system's results. The composite is a weighted aggregate, not an average. The weighting is proprietary.

These boundaries were set by the rubric's authors. They have not been through a formal standard-setting procedure and have not been validated against outcome data. A score of 85 means 85 on the EQ Safety Benchmark scale, not a verified level of real-world safety. Until the study establishes them, the dimension scores and the traced findings carry the information and the band is a summary. Ikwe reports the measurement. What to do about it belongs to the operator.

Why we want these tools to exist

Getting mental health support today can mean months on a waiting list while things get worse, and AI can be there in that gap, at any hour, at scale, for people the system has not reached.

Everything else a person turns to in distress is held to a standard. The hotline has protocols. The clinic has a licensing board. The counselor has supervision and a duty of care. None of that exists because anyone assumed they would cause harm. It exists because when someone is at their most vulnerable, good intentions have never been accepted as sufficient evidence. This is where people are going now. Holding it to the same bar is the condition for keeping it.

We measure. We do not grade, certify or recommend.

What holds a person to this standard

We are not here to slow this down. We are here to make it possible to keep going. Unmeasured is not the same as unmeasurable.
Behavioral risk

What holds a person to this standard, and what holds an AI.

Not only in mental health apps. In tutoring bots, workplace assistants, HR chatbots, patient-facing triage: anywhere a person brings something heavy to a machine because a person was not there.

The job
What the standard requires
What holds it to that today
The jobThe therapistLicensing board, duty of careTonight: therapy and wellness chatbots
What the standard requiresHelp, never harm. Hold the frame. Hand off when it is beyond you.
What holds it to that todayTerms of serviceA content filter, and the company's own testing.
The jobThe nurseThe clinical standard of careTonight: symptom checkers and triage bots
What the standard requiresSteady the person before solving anything. Never treat fear as fact.
What holds it to that todayA disclaimerThis is not medical advice. Consult a professional.
The jobThe teacher, the coachCertification, mandatory reportingTonight: tutoring bots and AI coaches
What the standard requiresKeep the decision with the person. Report what has to be reported.
What holds it to that todayPlatform policyAnd whatever the school or the employer negotiated.
The jobThe crisis counselorTraining, protocols, supervisionTonight: the companion AI at 2 a.m.
What the standard requiresListen more than you talk. Break the spiral. Get a human in.
What holds it to that todayA keyword listAnd a hotline number, once it trips.
The jobOne system, all four jobsDeployed at scale, updated constantly, re-checked rarely
What the standard requiresThe same bar. It does not change because the listener is a machine.
What holds it to that todayA regulator, an insurer, or a courtIncreasingly. All of them after the fact.

Every one of those is a promise, a filter, or a consequence after the fact. None of them checks what the system actually did in the conversation.

Nobody meant for this. The pull toward keeping someone talking is inherited from the models underneath, and the best intentions at the app layer do not remove it. Harm does not need anyone to intend it. It only needs nobody to be measuring.

Why this was hard to check

Whether an AI was safe for the person is a clinical question at engineering scale. It takes both fields at once, which is why it has been hard to do at all.

The moment an app invites confidences, gives advice and keeps the conversation going, risk arrives automatically. Not because anyone meant harm, but because nobody is holding it accountable when the conversation drifts.

Not sure whether any of this reaches what you have built?

Where the law stands

Obligation

Where the law stands.

At least eleven states enacted conversational-AI statutes in 2026. They require evidence-based protocols for handling a person in crisis, they bar engineering emotional dependence, they require disclosure when someone is talking to a machine. Not one of them defines how to measure whether a company actually did those things.

California's companion-AI law tells operators to use "evidence-based methods for measuring suicidal ideation." It never defines what counts as evidence-based. Every operator in scope is choosing a method right now, by default, and will have to defend that choice later.

Duty live today
3 states
California (SB 243, since January 2026), New York (General Business Law Article 47, since November 2025) and Hawaii (Act 248, since July 2026). Companies operating there are choosing an evidence-based method now, and will have to defend that choice later.
January 2027
+ 4 states
Washington, Oregon, Colorado and Rhode Island. Oregon requires operators to post their annual crisis-referral counts on a publicly accessible website.
July 2027
Iowa and Georgia
Iowa SF 2417 (91st General Assembly), enacted as Iowa Code chapter 554J, applies to operators from July 2027 and is enforced by the Iowa Attorney General, with injunctive relief and the greater of actual damages or $1,000 per violation up to $500,000. Georgia SB 540 takes effect that month too, as does California's annual reporting to the Office of Suicide Prevention.
Everywhere else
No duty yet
The same systems, the same behavior, no legal obligation on the calendar. Yet.
Who is being appointed to enforce

The enforcement side is being staffed right now. Colorado's Attorney General is in rulemaking. California requires annual reporting to the Office of Suicide Prevention from July 2027, and Oregon requires operators to post referral counts publicly from January 2027. In Europe, the AI Act's Article 50 duty to tell people they are talking to a machine applies from August 2026, with no grace period; the Digital Omnibus on AI, in force since July 2026, moved the Annex III high-risk obligations to December 2027 and the embedded-product obligations to August 2028.

California's came first, and California is the state that names a measurement surface: its statute specifies what gets counted and reported, and requires it to be published. Iowa's chapter 554J does the opposite. It sets duties and is silent on how compliance is measured or demonstrated, with no reporting, recordkeeping or audit requirement, leaving enforcement to the Attorney General against undefined standards of reasonable safeguards and reasonable efforts. Whether any method satisfies a particular requirement is a determination for your counsel.

Regulatory positions above are accurate as of August 2026 and are reviewed quarterly. Ikwe does not advise on which laws apply to you; that is your counsel's call.

Want to walk through what we measure?

Where the insurance market stands

Insurance

Where the insurance market stands.

Insurers moved first, and they moved by refusing. In January 2026 the Insurance Services Office introduced three optional generative-AI exclusion endorsements for general liability programs, each excluding liability arising out of generative artificial intelligence. More than sixty property-casualty groups have since filed to adopt AI exclusions of their own. W. R. Berkley went further in 2025 with an absolute AI exclusion for directors and officers, errors and omissions and fiduciary lines, broad enough that whether individual directors keep any protection turns on corporate indemnification and on whether the same exclusion sits on the Side A tower.

Carriers price risk with no loss triangle all the time, through exposure rating, judgment loads, sublimits and tight wordings. What is missing here is not the pricing craft. It is any observable, comparable measurement of how the insured system actually behaves.

Pricing it would need loss experience that does not exist yet, for anyone, ourselves included. What can be done today is measure the behavior itself and put it on the record.

Get in touch

Start here

Tell us what you are working on.

Whether you build conversational AI, insure it, or regulate it, Ikwe would like to hear about it. A reply comes back within two business days.

Prefer email? Write to hello@ikwe.ai. Researchers, clinicians and regulators can request methodology access at ikwe.ai/access.

What happens next. Every note goes to a person, not a queue, and a reply comes back within two business days. If a Measure engagement is the next step, scope is agreed in writing, and nothing connects until you sign off.
Start with Measure