Research
Our research is driven by advanced Computer Vision and Machine Learning & Robot Vision, equipping machines to truly understand complex real-world environments. Our research spans 3D Reconstruction for precise geometric scene understanding, Social Kinematics for modeling human-agent dynamics, and LLMs & Creative AI for generating physics-based content. Especially, our group's research achievements are listed in below:
Autonomous Driving
- Trajectory Prediction & Reasoning
- Motion & Behavior Generation
Related Works:
AAAI'21, CVPR'22, ECCV'22, AAAI'23ORAL, ICCV'23, CVPR'24(1), CVPR'24(2), CVPR'25, IEEE TPAMI'26, NeurIPS'26(Under-Review)
Physical AI
- Vision-Language Action Models
- Robotics & Physical Simulation
Related Works:
CVPR'24(1), NeurIPS'26(Under-Review)
Generative AI
- Video & Image Content Creation
- Sign Language Generation
- TEM / Medical Image Analysis
Related Works:
CVPR'24(2), ECCV'24, NeurIPS'26(Under-Review)
3D Reconstruction
- 2D, 3D, 4D Scene Reconstruction
- Depth Estimation & Completion
Related Works:
ICML'23, ICLR'24, NeurIPS'24, ICCV'25HIGHLIGHT, ECCV'26
Large Language Models
- Vision-Language Alignment
- Reasoning via VLMs
Related Works:
CVPR'24(2), IEEE TPAMI'26, NeurIPS'26(Under-Review)
Machine Learning for CV
- Machine Learning Theory
- Mathematical Formulation
Related Works:
CVPR'22, ICML'23, ICCV'23