This dissertation presents temporal-difference learning methods for building world models that enable intelligent agents and robots to predict, plan and act. It introduces TD-MPC, which learns compact latent dynamics, rewards and value estimates for model-predictive control, and TD-MPC2, a more robust successor requiring little task-specific tuning. Experiments demonstrate faster real-world robot learning, scalable multitask performance and a unified model spanning 200 tasks across ten domains. The research also examines interactive generative world models and their tendency to hallucinate physically implausible futures. Targeted data collection in uncertain regions substantially reduces these errors, suggesting routes toward safer, uncertainty-aware autonomous systems and robotics.

Babies are exceptional learners, possibly because they use surprise to guide attention and learning. My research shows that infants learn more after surprising physical or social events. Adults show a Goldilocks effect—optimal learning from moderate surprise. Understanding surprise-based learning in babies may help improve future artificial intelligence systems.