2026
This dissertation presents temporal-difference learning methods for building world models that enable intelligent agents and robots to predict, plan and act. It introduces TD-MPC, which learns compact latent dynamics, rewards and value estimates for…