This dissertation presents temporal-difference learning methods for building world models that enable intelligent agents and robots to predict, plan and act. It introduces TD-MPC, which learns compact latent dynamics, rewards and value estimates for model-predictive control, and TD-MPC2, a more robust successor requiring little task-specific tuning. Experiments demonstrate faster real-world robot learning, scalable multitask performance and a unified model spanning 200 tasks across ten domains. The research also examines interactive generative world models and their tendency to hallucinate physically implausible futures. Targeted data collection in uncertain regions substantially reduces these errors, suggesting routes toward safer, uncertainty-aware autonomous systems and robotics.

This thesis proposes scaling humanoid robots through large-scale motion imitation, learned control priors, and reinforcement learning. A perpetual humanoid controller achieves full-dataset imitation, distilled into a universal latent space enabling efficient learning of manipulation and vision-based tasks. The approach transfers from simulation to real robots, advancing practical humanoid control.