REINFORCE

E426681

REINFORCE is a classic Monte Carlo policy gradient algorithm in reinforcement learning that optimizes stochastic policies by estimating gradients from sampled returns.

All labels observed (5)

How this entity was disambiguated

Statements (47)

Predicate Object
instanceOf Monte Carlo reinforcement learning algorithm ⓘ
on-policy reinforcement learning method ⓘ
policy gradient algorithm ⓘ
applicableTo continuous action spaces ⓘ
discrete action spaces ⓘ
assumes differentiable policy with respect to parameters ⓘ
baselineType state-dependent baseline ⓘ
value function baseline ⓘ
canUse baseline to reduce variance ⓘ
category policy search method ⓘ
commonImplementation neural network policy ⓘ
creditAssignment returns assigned to actions in trajectory ⓘ
doesNotRequire environment model ⓘ
estimates policy gradient ⓘ
explorationMechanism inherent in stochastic policy ⓘ
field reinforcement learning ⓘ
gradientEstimator sampled returns ⓘ
gradientFormula E[ G_t ∇_θ log π_θ(a_t|s_t) ] ⓘ
influenced REINFORCE with baseline variants ⓘ
advantage actor-critic algorithms ⓘ
input trajectories of states, actions, rewards ⓘ
inspired later actor-critic methods ⓘ
introducedBy Ronald J. Williams ⓘ
introducedInPaper Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning ⓘ
linked to: REINFORCE
learningParadigm model-free reinforcement learning ⓘ
limitation high variance of gradient estimates ⓘ
sample inefficiency ⓘ
objective maximize expected cumulative reward ⓘ
optimizationMethod stochastic gradient ascent ⓘ
optimizes stochastic policies ⓘ
output updated policy parameters ⓘ
policyRepresentation parameterized function approximator ⓘ
policyType stochastic policy ⓘ
publicationYear 1992 ⓘ
relatedTo likelihood ratio gradient estimator ⓘ
score function estimator ⓘ
requires complete episodes ⓘ
differentiable log π_θ(a|s) ⓘ
strength conceptual simplicity ⓘ
does not require value function estimation ⓘ
trainingSignal sampled return from environment ⓘ
updateDirection proportional to return times log-probability gradient ⓘ
updateFrequency episode-wise updates ⓘ
updateRule gradient ascent on expected return ⓘ
usedIn episodic reinforcement learning settings ⓘ
uses Monte Carlo returns ⓘ
varianceProperty high variance gradient estimates ⓘ

How these facts were elicited

Referenced by (8)

Full triples — surface form annotated when it differs from this entity's canonical label.

TRPO → relatedTo → REINFORCE ⓘ
Ronald J. Williams → knownFor → REINFORCE algorithm ⓘ
linked to: REINFORCE
Ronald J. Williams → coAuthorOf → “Simple statistical gradient-following algorithms for connectionist reinforcement learning” ⓘ
linked to: REINFORCE
Ronald J. Williams → developed → REINFORCE learning rule ⓘ
linked to: REINFORCE
REINFORCE → introducedInPaper → Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning ⓘ
linked to: REINFORCE
Hindsight Policy Gradients → extends → REINFORCE algorithm ⓘ
linked to: REINFORCE
Tianshou → supportsAlgorithm → REINFORCE ⓘ