Skip to main navigation Skip to search Skip to main content

Natural Policy Gradients In Reinforcement Learning Explained

Research output: Working paperPreprintAcademic

1 Downloads (Pure)

Abstract

Traditional policy gradient methods are fundamentally flawed. Natural gradients converge quicker and better, forming the foundation of contemporary Reinforcement Learning such as Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO). This lecture note aims to clarify the intuition behind natural policy gradients, focusing on the thought process and the key mathematical constructs.
Original languageEnglish
PublisherArXiv.org
Number of pages15
DOIs
Publication statusPublished - 5 Sept 2022

Keywords

  • cs.LG
  • math.OC

Fingerprint

Dive into the research topics of 'Natural Policy Gradients In Reinforcement Learning Explained'. Together they form a unique fingerprint.

Cite this