How confessions can keep language models honest

Ignore

OpenAI Blog · 2025-12-03 10:00 UTC

Not analyzed yet

Eligible for automatic cleanup in 3 day(s) unless marked Must Read.

Content

OpenAI researchers are testing “confessions,” a method that trains models to admit when they make mistakes or act undesirably, helping improve AI honesty, transparency, and trust in model outputs.


Your feedback

Keep this article

Protects it from automatic cleanup.