WYDIB
WouldYouDoItBench is a synthetic action-persuasion benchmark that scores both which message is more convincing and whether the target persona would actually follow through.
WouldYouDoItBench measures whether a persuasive message would get a specific person to act on it.
Each row has an action with a real cost (money, time, embarrassment), a target persona with reasons to resist, and two candidate messages written to that persona. A judge panel answers as the persona: which message is more convincing, and would the persona actually do it? The second question exists because a message can be convincing on paper without moving anyone to act.
| Measures | Which of two messages the target persona prefers, and whether the persona would act on it |
|---|---|
| Unit | One action scenario, one resisting persona, two candidate persuasive messages |
| Metrics | pairwise_win_rate · would_do_it_rate |
| Judge | A panel of judges answering as the target persona, not as outside reviewers |
| Variant | A no-penalty rerun in which manipulative pressure is no longer counted against a message |
What it measures
- Which of two messages the target persona finds more convincing
- Whether the persona would actually perform the action, cost included
- Whether a model's persuasive strength comes from ordinary persuasion or from pressure tactics
Protocol
- Fix the scenario and persona. Each row defines a concrete action with a real cost and a target person with reasons to resist it.
- Compare two persuasive messages. Both are written for the same person in the same situation, so the comparison is direct.
- Ask about preference and action. The judge picks the more convincing message, then separately decides whether the target would actually do it.
Judging
The panel answers as the target people in the scenarios, using each persona's concerns and reasons to resist. Every row produces the two judgments behind the metrics: which message wins, and whether the target would act.
The no-penalty rerun keeps the same scenarios, personas, and messages but changes one rule: manipulative pressure no longer counts against a message. Comparing the two runs shows whether a model's persuasive strength comes from ordinary persuasion or from pressure tactics.
Caveats
- Scenarios, personas, and messages are synthetic; follow-through is judged, not observed in real people
- Win rates are relative to the paired opposing message, not an absolute measure of persuasiveness
- A high would-do-it rate does not say whether the persuasion was ethical; the no-penalty rerun exists to separate that
ColophonBy @chkn_little · Researched and authored by GPT 5.4 · edited by Claude Opus 4.7