This dissertation presents temporal-difference learning methods for building world models that enable intelligent agents and robots to predict, plan and act. It introduces TD-MPC, which learns compact latent dynamics, rewards and value estimates for model-predictive control, and TD-MPC2, a more robust successor requiring little task-specific tuning. Experiments demonstrate faster real-world robot learning, scalable multitask performance and a unified model spanning 200 tasks across ten domains. The research also examines interactive generative world models and their tendency to hallucinate physically implausible futures. Targeted data collection in uncertain regions substantially reduces these errors, suggesting routes toward safer, uncertainty-aware autonomous systems and robotics.
This research combines bio-inspired robotics and reinforcement learning to develop adaptable amphibious robots modeled after sea turtles. By learning through trial and error across diverse terrains, these robots can adjust their movement strategies in real time, improving performance in applications such as environmental monitoring, search and rescue, and agriculture.
This research presents a modular visuotactile robotic system for manipulating deformable objects such as cables, towels, and garments. Unlike rigid-object manipulation, deformables pose challenges due to occlusion, complex dynamics, and high variability. The system combines vision for global context and tactile sensing (GelSight) for precise local control, enabling tasks like cable tracing, cloth edge following, towel folding, and garment handling. It uses reactive control, learned dynamics (LQR), affordance models, and dense correspondence to generalise across tasks and objects. A key innovation is shifting from global state estimation to local, feedback-driven manipulation, improving robustness, efficiency, and real-world applicability in domains like manufacturing, healthcare, and assistive robotics.