PaperBench: Evaluating AI’s Ability to Replicate AI Research

Ignore

OpenAI Blog · 2025-04-02 10:15 UTC

Not analyzed yet

Eligible for automatic cleanup in 2 day(s) unless marked Must Read.

Content

We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.


Your feedback

Keep this article

Protects it from automatic cleanup.