|
Building general-purpose embodied agents requires moving beyond specialized systems toward models that can understand,
reason about, and interact with diverse physical environments. In this talk, Zeyu Zhang will present his recent research on general embodied reasoning, covering 3D scene understanding, long-horizon navigation, robotic control, domain adaptation, and self-evolving
agents. The talk will discuss how multimodal pretraining, scalable synthetic data, and data-centric learning can produce transferable representations that support a wide range of tasks, including captioning, grounding, question answering, dialogue, spatial
reasoning, and planning. It will then examine how environmental understanding can be extended to embodied decision-making through hierarchical frameworks that combine fast reactive control with slower long-horizon reasoning. Zeyu will also discuss how carefully
designed post-training, parameter-efficient adaptation, and verifiable feedback can improve navigation and control while preserving previously acquired perception and reasoning capabilities. Finally, the talk will explore how embodied agents can continue improving
when manually specified objectives are incomplete or inaccurate, using self-supervised feedback, world-model-based exploration, adaptive memory, and stage-aware learning signals. Together, these studies highlight the importance of scalable data, spatially
grounded foundation models, restrained reinforcement learning, and continual self-improvement for advancing toward general embodied intelligence.
|