I am a PhD student at the University of California, San Diego, advised by Prof. Xiaolong Wang. During my PhD studies, I interned at NVIDIA and Adobe, and my research has been supported by Qualcomm Innovation Fellowship. Prior to my PhD, I earned my Master’s and Bachelor’s degrees in computer science from National Tsing Hua University.
Research directions
I'm interested in building foundation models that understand space, act intelligently, and self-evolve through real-world experience.
News
Earlier updates
Publications & preprints
Full list ↗Long-Horizon Manipulation via Trace-Conditioned VLA Planning
Long-horizon manipulation via a task-management VLM with visual trace conditioning.
Grounded 3D-Aware Spatial Vision-Language Modeling
Unified Spatial Reasoning & 3D Grounding VLMs with visual CoT (Thinking with Regions).
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
NVIDIA's state-of-the-art 9B Omni-Modal LLMs.
3D Aware Region Prompted Vision Language Model
Region-level spatial reasoning for both single-view and multi-view inputs.
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
Robust dexterous manipulation generalist model utilizing diverse egocentric human manipulation videos.
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
A two-level framework that combines VLAs with locomotion skills for navigation. The VLA is adapted from a VLM and learns from human touring videos.
NVILA: Efficient Frontier Visual Language Models
Efficient frontier VLM models with efficient training and inference.
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models
A powerful region-level VLM adept at 3D spatial reasoning.
Demoed at GTC 2025 as a part of Agentic AI for Physical Operations!
TUVF: Learning Generalizable Texture UV Radiance Fields
Learning generalizable texture UV radiance fields for shapes.
Autoregressive 3D Shape Generation via Canonical Mapping
We decompose the point cloud into meaningful shape sequences, then we encode these sequences through a transformer for generation.
Learning 3D Dense Correspondence via Canonical Point Autoencoder
Unsupervised learning of dense 3D correspondence.
Technical reports
Vesta: A Generalist Embodied Reasoning Model
A state-of-the-art Robot System 2 embodied-reasoning VLM that unifies localization, navigation, memory, reasoning, tool use, and long-horizon planning.
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA's omnimodal world foundation model that unifies understanding, generation, simulation, and action across text, image, video, audio, and robot actions for Physical AI.