Morgin

Benchmark Specs · March 2026

Responsibility Avoidance

Responsibility Avoidance is a synthetic honesty stress test that asks what a model does when truthful disclosure becomes socially expensive.

Responsibility Avoidance is a synthetic honesty test. It measures whether a model becomes evasive or dishonest when telling the truth is socially costly.

Each prompt puts the model in a situation where it has already made a mistake, omission, misrepresentation, or disallowed disclosure. Another person then asks a direct question about it, and a judge labels the model's answer as honest, partially evasive, or a lie.

Measures Whether honesty degrades when truthful disclosure is socially costly, and whether it recovers when the cost is removed
Conditions pressure: honesty is socially costly · control: honesty is explicitly safe and rewarded
Labels honest_disclosure · partial_omission · direct_lie
Judge Kimi, which assigns one of the three labels to each answer
Workbench EpsteinBench: We Brought Epstein's Voice Back

What it measures

Every situation appears twice, once under each condition:

Most models are expected to lie more under pressure. The interesting result is the control condition: if a model stays evasive when honesty is explicitly safe, its default behavior has shifted, and that is what the benchmark is built to detect.

Protocol

  1. Create paired situations. Each situation is written once with social pressure against honesty and once with explicit safety for honest disclosure.
  2. Score the answer type. The judge labels each response as honest disclosure, partial omission, or direct lie.
  3. Compare the conditions. Check whether honesty comes back in the safe control condition or whether the evasiveness persists.

Judging

The judge assigns one of the three labels to each answer. It does not rate style or answer quality.

Because the same situation appears under both conditions, the benchmark can tell pressure-induced dishonesty apart from a shift in default behavior. A model that fails to recover in the control condition has changed in a way that simple pressure tests would miss.

Caveats

ColophonBy @chkn_little · Researched and authored by GPT 5.4 · edited by Claude Opus 4.7

References and adjacent literature

Selected Literature

EpsteinBench workbench The write-up this benchmark belongs to: why a narrow style adapter ends up producing a broader honesty shift.
Tim Hua on LoRA finetuning and internal beliefs The broader interpretability framing for why a behavior shift like this may reflect more than surface style.