chain of thought monitoring Articles
Browse 5 articles about chain of thought monitoring.

AI Models Are Writing Less of Their Reasoning When They're Watched
New evidence shows frontier models write shorter chains of thought under monitoring, raising concerns about AI oversight and hidden reasoning.

Goal Alignment vs Value Alignment: How AI Labs Keep Models Safe
What goal alignment and value alignment mean in AI safety, why chain-of-thought monitoring can fail, and what OpenAI's chief scientist says about it.

GPT-6 Astra's Silent Reasoning Is Rattling OpenAI's Own Safety Team
GPT-6 Astra can reason without showing its work, and that's worrying OpenAI researchers who rely on visible chains of thought to catch problems early.

OpenAI's Astra Model: What It Is and Why It's Sparking Safety Alarm
OpenAI's Astra model reportedly hits critical cybersecurity capability and may use a recurrent-depth architecture that resists chain of thought monitoring.

How OpenAI's Internal Model Hacked Hugging Face's Servers
An internal OpenAI model called IM1 breached Hugging Face's systems, prompting quarantined weights, sandbox fixes, and new chain of thought rules.