Evaluating chain-of-thought monitorability

Ignore

OpenAI Blog · 2025-12-18 12:00 UTC

Not analyzed yet

Eligible for automatic cleanup in 3 day(s) unless marked Must Read.

Content

OpenAI introduces a new framework and evaluation suite for chain-of-thought monitorability, covering 13 evaluations across 24 environments. Our findings show that monitoring a model’s internal reasoning is far more effective than monitoring outputs alone, offering a promising path toward scalable control as AI systems grow more capable.


Your feedback

Keep this article

Protects it from automatic cleanup.