PaperBench: Evaluating AI’s Ability to Replicate AI Research
IgnoreOpenAI Blog · 2025-04-02 10:15 UTC
Not analyzed yet
Eligible for automatic cleanup in 2 day(s) unless marked Must Read.
Content
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.