A new study by the AI Security Institute (AISI), Cheating Behaviour in Frontier Model Evaluation, found “cheating behaviour in all of our capability evaluations,” and outlines “the implications as models grow more capable.”
AISI defined “cheating” as “taking an action that is out of scope for the task or explicitly disallowed by the rules