Deliberative alignment: reasoning enables safer language models
IgnoreOpenAI Blog · 2024-12-20 10:00 UTC
Not analyzed yet
Eligible for automatic cleanup in 2 day(s) unless marked Must Read.
Content
Deliberative alignment: reasoning enables safer language models Introducing our new alignment strategy for o1 models, which are directly taught safety specifications and how to reason over them.