Detecting misbehavior in frontier reasoning models
IgnoreOpenAI Blog · 2025-03-10 10:00 UTC
Not analyzed yet
Eligible for automatic cleanup in 2 day(s) unless marked Must Read.
Content
Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.