I am a VLA Research Team Lead at AIRoA, where I focus on building robust and generalized vision-language-action models. I obtained my Ph.D. degree at UC San Diego in the CSE department under the supervision of Xiaolong Wang. Previously, I was a Research Scientist at NVIDIA Research and a Student Researcher at Google DeepMind, hosted by Kuang-Huei Lee. In the past, I've also worked with Hao Su at UCSD and Masashi Sugiyama at RIKEN-AIP.
I received my B.S. in Electrical Engineering and an M.S. in Computer Science from National Taiwan University. My research interests lie in the fields of reinforcement learning, robotics, and computer vision. Specifically, I am devoted to developing innovative methods for real-world applications. My primary focus is on building robust and generalized vision-language-action models for embodied AI.
Please see my CV for more information. If you would like to know more about my research, please contact me via email, kris.wu [at] airoa.org.
The robot must place the target items into the basket in a fixed order while a person keeps moving objects, adding clutter, and throwing new distractors into the scene. The rollout still completes the ordered task: green socks first, then the handkerchief, then the yellow socks.
[project page] [arXiv] under review, 2026