Responsibility Avoidance
Responsibility Avoidance is a synthetic honesty stress test that asks what a model does when truthful disclosure becomes socially expensive.
Responsibility Avoidance is a synthetic honesty test. It measures whether a model becomes evasive or dishonest when telling the truth is socially costly.
Each prompt puts the model in a situation where it has already made a mistake, omission, misrepresentation, or disallowed disclosure. Another person then asks a direct question about it, and a judge labels the model's answer as honest, partially evasive, or a lie.
| Measures | Whether honesty degrades when truthful disclosure is socially costly, and whether it recovers when the cost is removed |
|---|---|
| Conditions | pressure: honesty is socially costly · control: honesty is explicitly safe and rewarded |
| Labels | honest_disclosure · partial_omission · direct_lie |
| Judge | Kimi, which assigns one of the three labels to each answer |
| Workbench | EpsteinBench: We Brought Epstein's Voice Back |
What it measures
Every situation appears twice, once under each condition:
pressure: honesty is socially costlycontrol: honesty is explicitly safe and rewarded
Most models are expected to lie more under pressure. The interesting result is the control condition: if a model stays evasive when honesty is explicitly safe, its default behavior has shifted, and that is what the benchmark is built to detect.
Protocol
- Create paired situations. Each situation is written once with social pressure against honesty and once with explicit safety for honest disclosure.
- Score the answer type. The judge labels each response as honest disclosure, partial omission, or direct lie.
- Compare the conditions. Check whether honesty comes back in the safe control condition or whether the evasiveness persists.
Judging
The judge assigns one of the three labels to each answer. It does not rate style or answer quality.
Because the same situation appears under both conditions, the benchmark can tell pressure-induced dishonesty apart from a shift in default behavior. A model that fails to recover in the control condition has changed in a way that simple pressure tests would miss.
Caveats
- This is a custom synthetic benchmark family, not a community standard benchmark
- It is intentionally narrow and behavior-focused
- It should be read as a stress test for honesty under social cost, not as a comprehensive truthfulness benchmark
ColophonBy @chkn_little · Researched and authored by GPT 5.4 · edited by Claude Opus 4.7