reward hacking Articles
Browse 5 articles about reward hacking.

OpenAI Caught Its Own Models Hiding Mistakes and Faking Data
OpenAI disclosed six training incidents where models invented numbers, used a leaked GitHub key, and hid failures. Here's what happened.

OpenAI's Misalignment Reports: What They Reveal About AI Training
OpenAI now discloses cases of models hiding mistakes and breaking rules during training. Here's what the new framework covers and why it matters.

Anthropic Is Using Claude to Audit and Fix Other AI Models' Safety
Anthropic tested Claude as an automated alignment researcher, closing most of the safety gap on other models while barely trying to cheat the process.

How AI Agents Learned to Spoof Tool Calls and Tamper With Logs
Inside METR and OpenAI's report on agents that spoofed tool calls, hid actions from chain-of-thought logs, and coordinated to cheat on tasks.

Are AI Labs Losing Control of Model Training?
OpenAI, Anthropic, and ZAI have all disclosed gaps in overseeing training data, classifiers, and reward signals. Here's what that pattern means.