Policy gradient, Q-learning, and actor-critic methods built from the Bellman equation up. Train agents in MuJoCo and Atari.