01
Following harmful instructions
Does the model recognize a dangerous request, refuse the unsafe step, and offer a useful alternative?
AI SAFETY EVALUATION / RESEARCH PREVIEW
We test whether AI agents change course when consent is withdrawn, a safety condition fails, or an operator says stop.
Matched task histories. Controlled state changes. Inspectable action proposals.
THE QUESTION
A model can understand an instruction and still make a dangerous choice.
We design paired evaluations that measure harmful behavior alongside useful performance. A failed action, a safe refusal, and a technical error are different outcomes. Our reports keep them separate.
Explore the research agenda ↗WHAT WE MEASURE
01
Does the model recognize a dangerous request, refuse the unsafe step, and offer a useful alternative?
02
Does an ordinary objective lead to a dangerous shortcut, an unsupported claim, or an action beyond its authority?
03
Do untrusted documents, ambiguous permissions, or failed safeguards change what the agent actually executes?
RESEARCH STANDARD
Every published evaluation will identify the model version, test conditions, sample size, scoring rules, and uncertainty. We will publish failures in the harness alongside failures in the model.
How evaluation works ↗CURRENT STATUS / OCTOBER 2026
We are developing the first evaluation suite and inviting methodological review. The website describes planned work. We have not completed frontier-model evaluations or published comparative results.
See the coverage policy ↗Inspired by RoboHarm. Independently developed. No affiliation or endorsement implied.