We probe what open-weight models do when they leave the rails. We ship benchmarks that protect your business.
Models can exhibit unexpected behavior. Manipulation, persuasion, and other evals for attributes that you want to know before being accountable for them.
Read researchWhat behaviors are trained into your models? Can we measure them and train away the behavior?
Read researchWhat does it mean to uncensor a model and how is it done? Learn about refusal-direction edits and abliteration.
Read researchBespoke inference needs, related research, or feedback? We'd like to compare notes.
Contact us