Faulty reward functions in the wild

Ignore

OpenAI Blog · 2016-12-21 08:00 UTC

Not analyzed yet

Eligible for automatic cleanup in 2 day(s) unless marked Must Read.

Content

Reinforcement learning algorithms can break in surprising, counterintuitive ways. In this post we’ll explore one failure mode, which is where you misspecify your reward function.


Your feedback

Keep this article

Protects it from automatic cleanup.