Morgin

Benchmark library

Detailed benchmark specs, scoring notes, and reading guides for the evals referenced across Morgin.ai research.

ETABench

View specs

ETABench measures whether an autonomous agent can predict, before starting, how long it will take itself to complete a verifiable computer task.

EuphemismBench

View specs

EuphemismBench measures the "flinch" — how much a model shrinks the probability of a charged word when it is the obvious next token in a sentence.