Proximal Policy Optimization
IgnoreOpenAI Blog · 2017-07-20 07:00 UTC
Not analyzed yet
Eligible for automatic cleanup in 2 day(s) unless marked Must Read.
Content
We’re releasing a new class of reinforcement learning algorithms, Proximal Policy Optimization (PPO), which perform comparably or better than state-of-the-art approaches while being much simpler to implement and tune. PPO has become the default reinforcement learning algorithm at OpenAI because of its ease of use and good performance.