Hi, I am Chengxuan Qian (ι±ζΏη«), a CS PhD Student (starting Fall 2026) at University of California, Santa Barbara (UCSB), advised by Prof. Yao Qin. Prior to that, I'm grateful for the mentorship of Prof. Zhengzhong Tu (TAMU), Prof. Yue Zhao (USC), and Prof. Han Liu (Northwestern University) during my undergraduate years.
My research focuses on identifying and eliminating barriers to foundation models reaching superintelligence. Foundation models grow from large-scale, static, and idealized settings, yet the real world is dynamic and partially observable. Can they actively explore, interact with tools and environments, and connect with memory? Can they simulate world dynamics, imagine future states, and continually evolve from experience? I aim to teach machines to think like humans and explore frontiers beyond human reach. Specifically, my recent research focuses on:
- Multimodal LLMs Reasoning and Post-Training: leveraging reinforcement learning, agent harnesses, and data scaling to enhance long-term memory, complex tool use, human-AI interaction, and self-evolution over long-horizon trajectories.
- Spatial and Physical Intelligence: building grounded multimodal physical intelligence that can understand, reason, imagine, interact, and act under limited-data conditions.
- Generative World Modeling: leveraging generative models to reason about and imagine how actions change world states in partially observable real-world environments.
- Generalizable Real-World Intelligence: building agents that can transfer knowledge and skills across tasks, environments, and embodiments in healthcare, robotics, autonomous driving, and beyond.
Our Lab is actively looking for Postdocs/PhD/visiting students/interns, with more details available here; junior students interested in collaborating with us, especially those at UCSB, are welcome to email chengxuanqian[at]ucsb.edu with your CV.
π₯News
- 2026.08: π₯π₯ Announcing SystemPromptIndex, the largest system prompt library, indexing over 1,000 system prompts from 400+ products, and AISPA, the first assurance standard for AI system prompts. [Paper] [Website] [X] [LinkedIn] [Synced (ζΊε¨δΉεΏ)]
- 2026.07: π₯π₯ We released our survey, Progress Reward Modeling for Robotic Learning, building on our earlier work ProgressLM. It ranked #2 on Hugging Face Daily Papers. [Paper] [Code]
- 2026.07: βοΈβοΈ Attending ACL 2026 in person in San Diego from July 5-8 and presenting ProgressLM in Oral Session F. [Poster] [Oral Presentation] [Slides]
- 2026.06: βοΈβοΈ Attending CVPR 2026 and presenting my work DynCIM and fMRI-LM in person from June 2-6. See you in Denver!
- 2026.05: ππ Recognized as a CVPR 2026 Outstanding Reviewer, top 5% of 17,491 reviewers.
- 2026.05: ππ Our work ProgressLM has been selected for an ACL 2026 Oral (Top 3.3%) π₯
- 2026.04: ππ ProgressLM, our general reward model for embodied agents, has been accepted to the ACL 2026 Main Conference and the ICLR 2026 Workshop on World Models!
- 2026.04: ππ My first-author work DynCIM on cross-modal imbalance in multimodal foundation models has been accepted by CVPR 2026 Workshop on Cognitive Foundations for Multimodal Models!
- 2026.03: π₯π₯ Joined University of California, Santa Barbara (UCSB) as a CS PhD student.
- 2026.02: ππ We release What If Agents Could Imagine?, a study that breaks through the static perception barrier of VLMs via active generative world modeling.
- 2026.02: ππ Our work fMRI-LM on Medical Foundation Models has been accepted by CVPR 2026!
- 2026.01: ππ Three first/co-first author papers have been accepted by ICLR 2026!
- DecAlign: Aligning Cross-Modal Semantics for Multimodal Foundation Models
- AutoDrive-RΒ²: Towards Physical-Grounded Multimodal Reasoning for Autonomous Driving
- Video-STAR: Tool-Augmented Agentic RL for Thinking with Videos
- 2026.01: ππ We propose ProgressLM, which further investigates whether VLMs can acquire human-like, generalizable mental understanding and simulation in embodied scenarios from a single example, and serves as an early step toward building general-purpose reward models. See More: [Website] [Paper] [Code] [Model] [Dataset]
- 2025.12: βοΈβοΈ Attended NeurIPS 2025 in-person, See you in San Diego!
- 2025.11: ππ Our work LiMT, an unified multi-task liver image benchmark work, has been accepted by Journal of Biomedical and Health Informatics (JBHI)!
- 2025.10: ππ Our work DVP-MVS++, Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo, has been accepted by IEEE Transactions on Circuits and Systems for Video Technology (IEEE TCSVT)!
- 2025.10: ππ My first-author work on Medical Segmentation under sparse and noisy labeled annotations has been accepted by BIBM AIBH 2025!
- 2025.10: ππ We propose Video-STAR, a powerful Tool-Augmented Agentic RL approach for Thinking with Videos. On open-vocabulary action recognition benchmarks like K-400 and HMDB-51, our 3B VLM achieves nearly 40% accuracy improvement over base models!π₯
- 2025.09: ππ Our work HAIF-GS, Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic Scene, has been accepted by NeurIPS 2025!
- 2025.09: ππ We propose AutoDrive-RΒ², Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving. We're also honored that our work was featured by AutoDrive Heart (θͺε¨ι©Ύι©ΆδΉεΏ)!
- 2025.08: ππ Our work Re-Align has been accepted by EMNLP 2025 Main Conference!
- 2025.07: ππ Our work on Generalizable Medical Vision has been Accepted by IEEE Transactions on Medical Imaging.
- 2025.06: βοΈβοΈ Attended CVPR 2025 in-person, See you in Nashville!
- 2025.05: ππ Our work CLIMD has been Early Accepted by MICCAI 2025 (Top 9%).
- 2025.03: ππ Excited to propose my first-author work DecAlign, a novel cross-modal decoupling and alignment framwork for multimodal representation learning.
- 2025.02: βοΈβοΈ Attended AAAI 2025 in-person, See you in Philadelphia!
- 2024.11: ππ Excited to propose my first-author work DynCIM, a novel dynamic multimodal curriculum learning framework in addressing cross-modal competition and imbalances, which is now available as a preprint!
- 2024.10: ππ We propose FASS, a novel frequency domain-enhanced approach for Medical Image Segmentation under Low-Contrast environment.
π Selected Publications (For the full list, please see Google Scholar)

Adaptive Label Correction for Robust Medical Image Segmentation with Noisy Labels
Medical Segmentation Noisy Labels Robust Learning
BIBM AIBH 2025
Chengxuan Qian, Kai Han, Jianxia Ding, Chongwen Lyu, Zhenlong Yuan, Jun Chen, Zhe Liu.

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
Multimodal Alignment Foundation Model Interpretability
ICLR 2026
Chengxuan Qian, Shuo Xing, Shawn Li, Yue Zhao, Zhengzhong Tuβ .

DynCIM: Dynamic Curriculum for Imbalanced Multimodal Learning
Multimodal Competition Foundation Models Curriculum Learning
CVPR 2026 Workshop on Cognitive Foundations for Multimodal Models
Chengxuan Qian, Kai Han, Jiaxin Liu, Zhenlong Yuan, Zhengzhong Zhu, Jingchao Wang, Chongwen Lyu, Jun Chen, Zhe Liu.

ProgressLM: Towards Progress Reasoning in Vision-Language Models
Embodied Agents VLM Post-Training Reasoning
Website Paper Code Model Dataset
ACL 2026 (Oral, Top 3.3%)
ICLR 2026 Workshop on World Models
Jianshu Zhang*, Chengxuan Qian*, Haosen Sun, Haoran Lu, Dingcheng Wang, Letian Xue, Han Liu

Multimodal Reasoning Autonomous Driving Open-World Applications
Featured by AutoDrive Heart (θͺε¨ι©Ύι©ΆδΉεΏ)
ICLR 2026
Zhenlong Yuan*, Chengxuan Qian*, Jing Tang, Jinguo Luo, Rui Chen, Lei Sun, Xiangxiang Chu, Yujun Cai, Dapeng Zhang, Shuo Li.

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
Think with Videos Tool-Using Agent Multi-turn Agentic RL
Paper Code 3B Model 7B Model Dataset
ICLR 2026
Zhenlong Yuan*, Xiangyan Qu*, Chengxuan Qian*, Rui Chen, Jing Tang, Lei Sun, Xiangxiang Chu, Dapeng Zhang, Yiwei Wang, Yujun Cai, Shuo Li.

CLIMD: A Curriculum Learning Framework for Imbalanced Multimodal Diagnosis
Multimodal Diagnosis Curriculum Learning Class Imbalance
MICCAI 2025 (Early Accept, Top 9%)
Kai Han, Chongwen Lyu, Lele Ma, Chengxuan Qian, Siqi Ma, Zheng Pang, Jun Chen, Zhe Liu.

fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding
Medical Foundation Model fMRI Understanding Language Alignment
CVPR 2026
Yuxiang Wei, Yanteng Zhang, Xi Xiao, Chengxuan Qian, Tianyang Wang, Vince D. Calhoun.

DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo
Multi-View Stereo 3D Reconstruction Visibility Modeling
IEEE Transactions on Circuits and Systems for Video Technology, 2025
Zhenlong Yuan, Dapeng Zhang, Zehao Li, Chengxuan Qian, Jianing Chen, Yinda Chen, Kehua Chen, Tianlu Mao, Zhaoxin Li, Hao Jiang, Zhaoqi Wang.

HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic Scene
Gaussian Splatting Dynamic Scenes World Modeling
NeurIPS 2025
Jianing Chen, Zehao Li, Yujun Cai, Hao Jiang, Chengxuan Qian, Juyuan Kang, Shuqin Gao, Honglong Zhao, Tianlu Mao, Yucheng Zhang.

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
System Prompt Auditing AI Safety Human-Centered AI
Tech Repo 2026
Featured by Synced (ζΊε¨δΉεΏ)
Xiangning Lin*, Shenzhe Zhu*, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei.

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation
Multimodal LLMs Generative World Modeling Agentic RL
Preprint 2026
Zhenlong Yuan, Xiangyan Qu, Jing Tang, Rui Chen, Lei Sun, Ruidong Chen, Hongwei Yu, Chengxuan Qian, Xiangxiang Chu, Shuo Li, Yuyin Zhou.

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
EMNLP 2025 Main Conference
Multimodal Alignment Hallucination Mitigation Multimodal RAG DPO
Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tuβ .

Region Uncertainty Estimation for Medical Image Segmentation with Noisy Labels
Medical Segmentation Noisy Labels Uncertainty Estimation
IEEE Transactions on Medical Imaging, 2025
Kai Han, Shuhui Wang, Jun Chen, Chengxuan Qian, Chongwen Lyu, Siqi Ma, Chengjian Qiu, Victor S. Sheng, Qingming Huang, Zhe Liu.

Frequency Domain Unlocks New Perspectives for Abdominal Medical Image Segmentation
Medical Segmentation Low-contrast Environment Robustness
Preprint 2025
Kai Han, Siqi Ma, Chengxuan Qian, Jun Chen, Chongwen Lyu, Yuqing Song, Zhe Liu.

LiMT: A Multi-task Liver Image Benchmark Dataset
Medical AI Benchmarks Multi-task Unified Learning
Journal of Biomedical and Health Informatics (JBHI 2025)
Zhe Liuβ , Kai Han, Siqi Ma, Yan Zhu, Jun Chen, Chongwen Lyu, Xinyi Qiu, Chengxuan Qian, Yuqing Song, Yi Liu, Liyuan Tian, Yang Ji, Yuefeng Li

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey
Robotic Learning Progress Reward Modeling Survey
Tech Repo 2026
#2 on Hugging Face Daily Papers
Jianshu Zhang*, Keliang Wu*, Haoran Lu*, Anbang Liu, Ce Zhang, Weijie Yin, Chengxuan Qian, Xiyuan Yang, Zhenyu Pan, Guo Ye, Han Liu.
π Selected Honors & Awards
- ACL Oral Presentation (Top 3.3%), 2026
- CVPR Outstanding Reviewer (Top 5% of 17,491), 2026
- Fellowship, Department of Computer Science, UCSB, 2026
- Outstanding Graduate Award (Top 6.6%), Jiangsu University, 2026
- Research Rising Star Award ($2500), Arcadia University, 2025
- Excellent Academic Scholarship ($48,000), Arcadia University, 2024
- National-Level College Studentsβ Innovation and Entrepreneurship Funded Program (Β₯8000), 2023
- First-class Academic Scholarship (Top 3%), Jiangsu University (Β₯3000), 2023, 2024
ποΈ Invited Talks
π Academic Services
- Journal Reviewer: IEEE Transactions on Circuits and Systems for Video Technology (TCSVT), IEEE Transactions on Multimedia (TMM), Pattern Recognition (PR), IEEE Transactions on Robot Learning (T-RL), ACM Computing Surveys (CSUR), IEEE Journal of Biomedical and Health Informatics (JBHI), IEEE Open Journal of Engineering in Medicine and Biology (OJEMB).
- Conference Reviewer: ICME 2025-2026, AAAI 2026-2027, ICASSP 2026, CVPR 2026, NeurIPS 2026, ICLR 2027, ACL Rolling Review (ARR).
- Workshop Reviewer: ACL 2025 SRW, NeurIPS 2025 Imageomics, NeurIPS 2025 Efficient Reasoning, ICLR 2026 Workshop on Lifelong Agents, ICLR 2026 Workshop World Models, COLM 2026 Workshop on Lifelong Agents, COLM 2026 Workshop Efficient Reasoning, NeurIPS 2026 Workshop on Physical Understanding for Decision-Making.
π Education
University of California, Santa Barbara (UCSB)
PhD Student in Computer Science
Jiangsu University
Bachelor of Science, Mathematics and Applied Mathematics
πΌ Experiences
Northwestern University
Research AssistantAdvisor: Prof. Han Liu
Focus: Multimodal LLMs Post-Training and Embodied Agents.
Work: ProgressLM (ACL 2026 Oral); Progress Reward Modeling (Tech Repo)
Alibaba AMAP-ML
Research CollaboratorFocus: Multimodal LLMs Post-Training and Agentic AI.
Work: Video-STAR (ICLR 2026); AutoDrive-RΒ² (ICLR 2026); ImagineAgent
Texas A&M University
Research AssistantAdvisors: Prof. Zhengzhong Tu and Prof. Yue Zhao (USC)
Focus: Reasoning and Alignment for Multimodal Foundation Models.
Jiangsu University
Research AssistantFocus: Multimodal Representation Learning and Medical Foundation Models.
Work: DynCIM (CVPRW 2026); ALC (BIBM AIBH 2025); FASS; CLIMD (MICCAI 2025); RUE (IEEE TMI 2025); LiMT (IEEE JBHI 2025); HAIF-GS (NeurIPS 2025); DVP-MVS++ (IEEE TCSVT 2025); fMRI-LM (CVPR 2026)