Wenli Xiao

Research

Co-Lead

Self-Improving Vision-Language-Action Models with Data Generation via Residual RL

ICLR 2026

PLD (Probe, Learn, Distill) is a plug-and-play recipe for Vision-Language-Action (VLA) post-training. It is model agnostic, supporting both autoregressive and diffusion architectures, and can push success rates to 99%.

Wenli Xiao*, Haotian Lin*, Andy Peng, Haoru Xue, Tairan He, Yuqi Xie, Fengyuan Hu, Jimmy Wu, Zhengyi Luo, Linxi "Jim" Fan†, Guanya Shi, Yuke Zhu†

Co-Author

ASPIRE: Agentic Skills Discovery for Robotics

arXiv 2026

ASPIRE enables coding agents to discover reusable robot skills by inspecting execution feedback, debugging control programs, and evolving a growing skill library. Skills learned in simulation transfer to unseen tasks and real robots, reducing the cost of robot programming.

Runyu Lu*†, Yubo Wu*, Ethan Kou*, Max Fu, Wenli Xiao, Ajay Mandlekar, Yinzhen Xu, Guanya Shi, Ken Goldberg, Ang Chen, Mosharaf Chowdhury, Yuke Zhu†, Linxi "Jim" Fan†, Guanzhi Wang†

Co-Lead1 / 16

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

CoRL 2026

Physical Autoresearch on real-world Robot Fleet. ENPIRE lets coding agents autonomously improve robot manipulation policies through a closed-loop physical feedback system—automatic environment reset and verification, parallel robot rollouts, and evolutionary refinement—reaching a 99% success rate on challenging dexterous manipulation tasks.

Wenli Xiao*, Jia Xie*, Tonghe Zhang*, Haotian Lin*, Letian "Max" Fu, Haoru Xue, Jalen Lu, Yi Yang, Cunxi Dai, Zi Wang, Jimmy Wu, Guanzhi Wang, S. Shankar Sastry, Ken Goldberg, Linxi "Jim" Fan‡, Yuke Zhu‡, Guanya Shi‡

swipe to browse