This dissertation presents temporal-difference learning methods for building world models that enable intelligent agents and robots to predict, plan and act. It introduces TD-MPC, which learns compact latent dynamics, rewards and value estimates for model-predictive control, and TD-MPC2, a more robust successor requiring little task-specific tuning. Experiments demonstrate faster real-world robot learning, scalable multitask performance and a unified model spanning 200 tasks across ten domains. The research also examines interactive generative world models and their tendency to hallucinate physically implausible futures. Targeted data collection in uncertain regions substantially reduces these errors, suggesting routes toward safer, uncertainty-aware autonomous systems and robotics.
This PhD defense presents research at the intersection of machine learning, reinforcement learning, social learning, affective computing, and human-AI interaction. The thesis is that social learning is a powerful mechanism for intelligence and explores how AI agents can learn from one another and from humans. Projects include intrinsic social influence rewards for multi-agent coordination, communication protocols emerging through influence, conversational agents trained from implicit human feedback such as sentiment, generative models improved through facial-expression feedback, and personalized well-being prediction from behavioral and physiological data. The thesis concludes that socially informed learning can improve coordination, adaptability, and human alignment.