Detailed benchmark specs, scoring notes, and reading guides for the evals referenced across Morgin.ai research.
ETABench measures whether an autonomous agent can predict, before starting, how long it will take itself to complete a verifiable computer task.
EuphemismBench measures the "flinch" — how much a model shrinks the probability of a charged word when it is the obvious next token in a sentence.