Hi folks,
Abstract: Evaluating general-purpose robot policies rigorously remains one of robotics' hardest challenges, given that real-world evaluation is slow and costly while existing simulation benchmarks suffer from high setup overhead and a persistent sim2real gap. In this talk, we discuss how to rigorously evaluate real-world generalist robot policies at scale, covering the key pitfalls in current benchmarking practice and our approach to addressing them. We introduce principles for future-proofing evaluation, including embodiment-agnostic task design, scalable task generation, and diagnostic analysis tools, which are needed to build evaluations that keep pace with increasingly capable models.
Bio: Xuning Yang is a Senior Research Scientist at NVIDIA's Seattle Research Lab. Her research interests include robot foundation models, particularly on evaluation methods, and generalization capabilities, as well as a broad range of application scenarios from field robotics, indoor navigation, and robot manipulation. Prior to NVIDIA, she received her Ph.D. in Robotics from Carnegie Mellon University,