MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Ignore

OpenAI Blog · 2024-10-10 10:00 UTC

Not analyzed yet

Eligible for automatic cleanup in 2 day(s) unless marked Must Read.

Content

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.


Your feedback

Keep this article

Protects it from automatic cleanup.