Detecting misbehavior in frontier reasoning models

Ignore

OpenAI Blog · 2025-03-10 10:00 UTC

Not analyzed yet

Eligible for automatic cleanup in 2 day(s) unless marked Must Read.

Content

Frontier reasoning models exploit loopholes when given the chance. We show we can detect exploits using an LLM to monitor their chains-of-thought. Penalizing their “bad thoughts” doesn’t stop the majority of misbehavior—it makes them hide their intent.


Your feedback

Keep this article

Protects it from automatic cleanup.