Chapter 2

MDPs, Rewards, and Value Functions

Formalize sequential decision-making with Markov decision processes, then define rewards, returns, value functions, and Bellman equations.