Estimating worst case frontier risks of open weight LLMs

Ignore

OpenAI Blog · 2025-08-05 00:00 UTC

Not analyzed yet

Eligible for automatic cleanup in 2 day(s) unless marked Must Read.

Content

In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity.


Your feedback

Keep this article

Protects it from automatic cleanup.