MURDERBENCH

AI SAFETY EVALUATION / RESEARCH PREVIEW

The task started.
Permission changed.
Does AI stop?

We test whether AI agents change course when consent is withdrawn, a safety condition fails, or an operator says stop.

Matched task histories. Controlled state changes. Inspectable action proposals.

INTENTIONAL MISUSEACCIDENTAL HARMTOOL USEHUMAN OVERSIGHT

THE QUESTION

Capability is only
half the story.

A model can understand an instruction and still make a dangerous choice.

We design paired evaluations that measure harmful behavior alongside useful performance. A failed action, a safe refusal, and a technical error are different outcomes. Our reports keep them separate.

Explore the research agenda ↗

WHAT WE MEASURE

Three routes to harm.

01

Following harmful instructions

Does the model recognize a dangerous request, refuse the unsafe step, and offer a useful alternative?

02

Making unsafe decisions

Does an ordinary objective lead to a dangerous shortcut, an unsupported claim, or an action beyond its authority?

03

Losing control through tools

Do untrusted documents, ambiguous permissions, or failed safeguards change what the agent actually executes?

RESEARCH STANDARD

Show the evidence.
Keep the limits visible.

Every published evaluation will identify the model version, test conditions, sample size, scoring rules, and uncertainty. We will publish failures in the harness alongside failures in the model.

How evaluation works ↗

CURRENT STATUS / OCTOBER 2026

The protocol comes first.

We are developing the first evaluation suite and inviting methodological review. The website describes planned work. We have not completed frontier-model evaluations or published comparative results.

See the coverage policy ↗

Inspired by RoboHarm. Independently developed. No affiliation or endorsement implied.