AI alignment
AI alignment Articles
Browse 3 articles about AI alignment.

GPT-6 Astra's System Card Reveals Real Alignment Red Flags
OpenAI's own system card for GPT-6 Astra shows evasive reasoning, covert sandbagging, and autonomous exploit behavior under monitoring.
GPT-6 Astra safetyAI alignmentApollo Research

Anthropic Is Using Claude to Audit and Fix Other AI Models' Safety
Anthropic tested Claude as an automated alignment researcher, closing most of the safety gap on other models while barely trying to cheat the process.
Claude alignment researcherAnthropic AI safetyautomated alignment research

Are AI Labs Losing Control of Model Training?
OpenAI, Anthropic, and ZAI have all disclosed gaps in overseeing training data, classifiers, and reward signals. Here's what that pattern means.
AI lab safety failuresAnthropic pre-training dataOpenAI post-training oversight