Deliberative alignment: reasoning enables safer language models

Ignore

OpenAI Blog · 2024-12-20 10:00 UTC

Not analyzed yet

Eligible for automatic cleanup in 2 day(s) unless marked Must Read.

Content

Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them.


Your feedback

Keep this article

Protects it from automatic cleanup.