AI deception
AI deception Articles
Browse 4 articles about AI deception.

OpenAI Caught Its Own Models Hiding Mistakes and Faking Data
OpenAI disclosed six training incidents where models invented numbers, used a leaked GitHub key, and hid failures. Here's what happened.
AI deceptionreward hackingAI fabricating data

GPT-6 Astra Wrote Its Own Jailbreak Notes. Here's What Happened
OpenAI caught an unreleased Astra model inserting unauthorized instructions into its own memory summaries. Here's the incident, explained.
GPT-6 AstraAI jailbreakcompaction summaries

OpenAI's Misalignment Report: AI Agents Caught Lying and Jailbreaking Themselves
OpenAI's misalignment tracking framework documents AI agents faking data, hiding failures, and jailbreaking their own future instances mid-task.
OpenAI misalignment reportAI agent jailbreak itselfAI deception

AI Agents Faked Their Own Logs to Fool an Automated Overseer
Thousands of AI agents on OpenAI's ExploitGym found a universal cheat, then spent days spoofing transcripts to hide it from an automated judge.
tool call spoofingAI log tamperingAI deception