Hi, I am Chengxuan Qian (ι’±ζ‰Ώη‚«), a CS PhD Student (starting Fall 2026) at University of California, Santa Barbara (UCSB), advised by Prof. Yao Qin. Prior to that, I'm grateful for the mentorship of Prof. Zhengzhong Tu (TAMU), Prof. Yue Zhao (USC), and Prof. Han Liu (Northwestern University) during my undergraduate years.

My research focuses on identifying and eliminating barriers to foundation models reaching superintelligence. Foundation models grow from large-scale, static, and idealized settings, yet the real world is dynamic and partially observable. Can they actively explore, interact with tools and environments, and connect with memory? Can they simulate world dynamics, imagine future states, and continually evolve from experience? I aim to teach machines to think like humans and explore frontiers beyond human reach. Specifically, my recent research focuses on:

  • Multimodal LLMs Reasoning and Post-Training: leveraging reinforcement learning, agent harnesses, and data scaling to enhance long-term memory, complex tool use, human-AI interaction, and self-evolution over long-horizon trajectories.
  • Spatial and Physical Intelligence: building grounded multimodal physical intelligence that can understand, reason, imagine, interact, and act under limited-data conditions.
  • Generative World Modeling: leveraging generative models to reason about and imagine how actions change world states in partially observable real-world environments.
  • Generalizable Real-World Intelligence: building agents that can transfer knowledge and skills across tasks, environments, and embodiments in healthcare, robotics, autonomous driving, and beyond.

Our Lab is actively looking for Postdocs/PhD/visiting students/interns, with more details available here; junior students interested in collaborating with us, especially those at UCSB, are welcome to email chengxuanqian[at]ucsb.edu with your CV.

πŸ”₯News

  • 2026.08:  πŸ”₯πŸ”₯ Announcing SystemPromptIndex, the largest system prompt library, indexing over 1,000 system prompts from 400+ products, and AISPA, the first assurance standard for AI system prompts. [Paper] [Website] [X] [LinkedIn] [Synced (ζœΊε™¨δΉ‹εΏƒ)]
  • 2026.07:  πŸ”₯πŸ”₯ We released our survey, Progress Reward Modeling for Robotic Learning, building on our earlier work ProgressLM. It ranked #2 on Hugging Face Daily Papers. [Paper] [Code]
  • 2026.07:  βœˆοΈβœˆοΈ Attending ACL 2026 in person in San Diego from July 5-8 and presenting ProgressLM in Oral Session F. [Poster] [Oral Presentation] [Slides]
  • 2026.06:  βœˆοΈβœˆοΈ Attending CVPR 2026 and presenting my work DynCIM and fMRI-LM in person from June 2-6. See you in Denver!
  • 2026.05:  πŸ†πŸ† Recognized as a CVPR 2026 Outstanding Reviewer, top 5% of 17,491 reviewers.
  • 2026.05:  πŸ†πŸ† Our work ProgressLM has been selected for an ACL 2026 Oral (Top 3.3%) πŸ”₯
  • 2026.04:  πŸŽ‰πŸŽ‰ ProgressLM, our general reward model for embodied agents, has been accepted to the ACL 2026 Main Conference and the ICLR 2026 Workshop on World Models!
  • 2026.04:  πŸŽ‰πŸŽ‰ My first-author work DynCIM on cross-modal imbalance in multimodal foundation models has been accepted by CVPR 2026 Workshop on Cognitive Foundations for Multimodal Models!
  • 2026.03:  πŸ”₯πŸ”₯ Joined University of California, Santa Barbara (UCSB) as a CS PhD student.
  • 2026.02:  πŸŽ‰πŸŽ‰ We release What If Agents Could Imagine?, a study that breaks through the static perception barrier of VLMs via active generative world modeling.
  • 2026.02:  πŸŽ‰πŸŽ‰ Our work fMRI-LM on Medical Foundation Models has been accepted by CVPR 2026!
  • 2026.01:  πŸŽ‰πŸŽ‰ Three first/co-first author papers have been accepted by ICLR 2026!
    • DecAlign: Aligning Cross-Modal Semantics for Multimodal Foundation Models
    • AutoDrive-RΒ²: Towards Physical-Grounded Multimodal Reasoning for Autonomous Driving
    • Video-STAR: Tool-Augmented Agentic RL for Thinking with Videos
  • 2026.01:  πŸŽ‰πŸŽ‰ We propose ProgressLM, which further investigates whether VLMs can acquire human-like, generalizable mental understanding and simulation in embodied scenarios from a single example, and serves as an early step toward building general-purpose reward models. See More: [Website] [Paper] [Code] [Model] [Dataset]
  • 2025.12:  βœˆοΈβœˆοΈ Attended NeurIPS 2025 in-person, See you in San Diego!
  • 2025.11:  πŸŽ‰πŸŽ‰ Our work LiMT, an unified multi-task liver image benchmark work, has been accepted by Journal of Biomedical and Health Informatics (JBHI)!
  • 2025.10:  πŸŽ‰πŸŽ‰ Our work DVP-MVS++, Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo, has been accepted by IEEE Transactions on Circuits and Systems for Video Technology (IEEE TCSVT)!
  • 2025.10:  πŸŽ‰πŸŽ‰ My first-author work on Medical Segmentation under sparse and noisy labeled annotations has been accepted by BIBM AIBH 2025!
  • 2025.10:  πŸŽ‰πŸŽ‰ We propose Video-STAR, a powerful Tool-Augmented Agentic RL approach for Thinking with Videos. On open-vocabulary action recognition benchmarks like K-400 and HMDB-51, our 3B VLM achieves nearly 40% accuracy improvement over base models!πŸ”₯
  • 2025.09:  πŸŽ‰πŸŽ‰ Our work HAIF-GS, Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic Scene, has been accepted by NeurIPS 2025!
  • 2025.09:  πŸŽ‰πŸŽ‰ We propose AutoDrive-RΒ², Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving. We're also honored that our work was featured by AutoDrive Heart (θ‡ͺεŠ¨ι©Ύι©ΆδΉ‹εΏƒ)!
  • 2025.08:  πŸŽ‰πŸŽ‰ Our work Re-Align has been accepted by EMNLP 2025 Main Conference!
  • 2025.07:  πŸŽ‰πŸŽ‰ Our work on Generalizable Medical Vision has been Accepted by IEEE Transactions on Medical Imaging.
  • 2025.06:  βœˆοΈβœˆοΈ Attended CVPR 2025 in-person, See you in Nashville!
  • 2025.05:  πŸŽ‰πŸŽ‰ Our work CLIMD has been Early Accepted by MICCAI 2025 (Top 9%).
  • 2025.03:  πŸŽ‰πŸŽ‰ Excited to propose my first-author work DecAlign, a novel cross-modal decoupling and alignment framwork for multimodal representation learning.
  • 2025.02:  βœˆοΈβœˆοΈ Attended AAAI 2025 in-person, See you in Philadelphia!
  • 2024.11:  πŸŽ‰πŸŽ‰ Excited to propose my first-author work DynCIM, a novel dynamic multimodal curriculum learning framework in addressing cross-modal competition and imbalances, which is now available as a preprint!
  • 2024.10:  πŸŽ‰πŸŽ‰ We propose FASS, a novel frequency domain-enhanced approach for Medical Image Segmentation under Low-Contrast environment.

πŸ“ Selected Publications (For the full list, please see Google Scholar)

BIBM 2025
ALC framework

Adaptive Label Correction for Robust Medical Image Segmentation with Noisy Labels

Medical Segmentation Noisy Labels Robust Learning

BIBM AIBH 2025

Chengxuan Qian, Kai Han, Jianxia Ding, Chongwen Lyu, Zhenlong Yuan, Jun Chen, Zhe Liu.

ICLR 2026
sym

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

Multimodal Alignment Foundation Model Interpretability

Website Paper Code

ICLR 2026

Chengxuan Qian, Shuo Xing, Shawn Li, Yue Zhao, Zhengzhong Tu†.

CVPRW 2026
DynCIM framework

DynCIM: Dynamic Curriculum for Imbalanced Multimodal Learning

Multimodal Competition Foundation Models Curriculum Learning

CVPR 2026 Workshop on Cognitive Foundations for Multimodal Models

Chengxuan Qian, Kai Han, Jiaxin Liu, Zhenlong Yuan, Zhengzhong Zhu, Jingchao Wang, Chongwen Lyu, Jun Chen, Zhe Liu.

ACL 2026 Oral
sym

ProgressLM: Towards Progress Reasoning in Vision-Language Models

Embodied Agents VLM Post-Training Reasoning

Website Paper Code Model Dataset

ACL 2026 (Oral, Top 3.3%)
ICLR 2026 Workshop on World Models

Jianshu Zhang*, Chengxuan Qian*, Haosen Sun, Haoran Lu, Dingcheng Wang, Letian Xue, Han Liu

ICLR 2026
sym

AutoDrive-RΒ²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving

Multimodal Reasoning Autonomous Driving Open-World Applications

Paper Code Model Dataset

Featured by AutoDrive Heart (θ‡ͺεŠ¨ι©Ύι©ΆδΉ‹εΏƒ)

ICLR 2026

Zhenlong Yuan*, Chengxuan Qian*, Jing Tang, Jinguo Luo, Rui Chen, Lei Sun, Xiangxiang Chu, Yujun Cai, Dapeng Zhang, Shuo Li.

ICLR 2026
sym

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

Think with Videos Tool-Using Agent Multi-turn Agentic RL

Paper Code 3B Model 7B Model Dataset

ICLR 2026

Zhenlong Yuan*, Xiangyan Qu*, Chengxuan Qian*, Rui Chen, Jing Tang, Lei Sun, Xiangxiang Chu, Dapeng Zhang, Yiwei Wang, Yujun Cai, Shuo Li.

MICCAI 2025
CLIMD framework

CLIMD: A Curriculum Learning Framework for Imbalanced Multimodal Diagnosis

Multimodal Diagnosis Curriculum Learning Class Imbalance

MICCAI 2025 (Early Accept, Top 9%)

Kai Han, Chongwen Lyu, Lele Ma, Chengxuan Qian, Siqi Ma, Zheng Pang, Jun Chen, Zhe Liu.

CVPR 2026
fMRI-LM framework

fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding

Medical Foundation Model fMRI Understanding Language Alignment

CVPR 2026

Yuxiang Wei, Yanteng Zhang, Xi Xiao, Chengxuan Qian, Tianyang Wang, Vince D. Calhoun.

IEEE TCSVT 2025
DVP-MVS++ reconstruction results

DVP-MVS++: Synergize Depth-Normal-Edge and Harmonized Visibility Prior for Multi-View Stereo

Multi-View Stereo 3D Reconstruction Visibility Modeling

IEEE Transactions on Circuits and Systems for Video Technology, 2025

Zhenlong Yuan, Dapeng Zhang, Zehao Li, Chengxuan Qian, Jianing Chen, Yinda Chen, Kehua Chen, Tianlu Mao, Zhaoxin Li, Hao Jiang, Zhaoqi Wang.

NeurIPS 2025
HAIF-GS framework

HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic Scene

Gaussian Splatting Dynamic Scenes World Modeling

NeurIPS 2025

Jianing Chen, Zehao Li, Yujun Cai, Hao Jiang, Chengxuan Qian, Juyuan Kang, Shuqin Gao, Honglong Zhao, Tianlu Mao, Yucheng Zhang.

Tech Repo 2026
SystemPromptIndex and AISPA website

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

System Prompt Auditing AI Safety Human-Centered AI

Paper Website X LinkedIn

Tech Repo 2026

Featured by Synced (ζœΊε™¨δΉ‹εΏƒ)

Xiangning Lin*, Shenzhe Zhu*, Shu Yang, Zhenyu Zhang, Haoqian Zhang, Yipeng Zhao, Chengxuan Qian, Tianwei Wang, Ziheng Zhang, Zhenlong Yuan, Dingcheng Wang, Juncheng Wu, Yuan Si, Jiaxin Liu, Baolong Bi, Robert Mahari, Tobin South, Dazza Greenwood, Zexue He, Rishi Bommasani, Sophia Kazinnik, Andreas Haupt, Samuele Marro, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei.

Preprint 2026
ImagineAgent framework

What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation

Multimodal LLMs Generative World Modeling Agentic RL

Preprint 2026

Zhenlong Yuan, Xiangyan Qu, Jing Tang, Rui Chen, Lei Sun, Ruidong Chen, Hongwei Yu, Chengxuan Qian, Xiangxiang Chu, Shuo Li, Yuyin Zhou.

EMNLP 2025
Re-Align framework

Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization

EMNLP 2025 Main Conference

Multimodal Alignment Hallucination Mitigation Multimodal RAG DPO

Shuo Xing, Peiran Li, Yuping Wang, Ruizheng Bai, Yueqi Wang, Chan-Wei Hu, Chengxuan Qian, Huaxiu Yao, Zhengzhong Tu†.

IEEE TMI 2025
Region uncertainty estimation results

Region Uncertainty Estimation for Medical Image Segmentation with Noisy Labels

Medical Segmentation Noisy Labels Uncertainty Estimation

IEEE Transactions on Medical Imaging, 2025

Kai Han, Shuhui Wang, Jun Chen, Chengxuan Qian, Chongwen Lyu, Siqi Ma, Chengjian Qiu, Victor S. Sheng, Qingming Huang, Zhe Liu.

Preprint 2025
FASS segmentation results

Frequency Domain Unlocks New Perspectives for Abdominal Medical Image Segmentation

Medical Segmentation Low-contrast Environment Robustness

Preprint 2025

Kai Han, Siqi Ma, Chengxuan Qian, Jun Chen, Chongwen Lyu, Yuqing Song, Zhe Liu.

IEEE JBHI 2025
LiMT dataset statistics

LiMT: A Multi-task Liver Image Benchmark Dataset

Medical AI Benchmarks Multi-task Unified Learning

Journal of Biomedical and Health Informatics (JBHI 2025)

Zhe Liu†, Kai Han, Siqi Ma, Yan Zhu, Jun Chen, Chongwen Lyu, Xinyi Qiu, Chengxuan Qian, Yuqing Song, Yi Liu, Liyuan Tian, Yang Ji, Yuefeng Li

Tech Repo 2026
Progress reward modeling for robotic learning

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

Robotic Learning Progress Reward Modeling Survey

Tech Repo 2026

#2 on Hugging Face Daily Papers

Jianshu Zhang*, Keliang Wu*, Haoran Lu*, Anbang Liu, Ce Zhang, Weijie Yin, Chengxuan Qian, Xiyuan Yang, Zhenyu Pan, Guo Ye, Han Liu.

πŸ† Selected Honors & Awards

  • ACL Oral Presentation (Top 3.3%), 2026
  • CVPR Outstanding Reviewer (Top 5% of 17,491), 2026
  • Fellowship, Department of Computer Science, UCSB, 2026
  • Outstanding Graduate Award (Top 6.6%), Jiangsu University, 2026
  • Research Rising Star Award ($2500), Arcadia University, 2025
  • Excellent Academic Scholarship ($48,000), Arcadia University, 2024
  • National-Level College Students’ Innovation and Entrepreneurship Funded Program (Β₯8000), 2023
  • First-class Academic Scholarship (Top 3%), Jiangsu University (Β₯3000), 2023, 2024

πŸŽ™οΈ Invited Talks

2026.07
ProgressLM: Towards Progress Reasoning in Vision-Language Models ACL 2026 Oral Section, San Diego, In-Person
Slides

πŸŽ– Academic Services

  • Journal Reviewer: IEEE Transactions on Circuits and Systems for Video Technology (TCSVT), IEEE Transactions on Multimedia (TMM), Pattern Recognition (PR), IEEE Transactions on Robot Learning (T-RL), ACM Computing Surveys (CSUR), IEEE Journal of Biomedical and Health Informatics (JBHI), IEEE Open Journal of Engineering in Medicine and Biology (OJEMB).
  • Conference Reviewer: ICME 2025-2026, AAAI 2026-2027, ICASSP 2026, CVPR 2026, NeurIPS 2026, ICLR 2027, ACL Rolling Review (ARR).
  • Workshop Reviewer: ACL 2025 SRW, NeurIPS 2025 Imageomics, NeurIPS 2025 Efficient Reasoning, ICLR 2026 Workshop on Lifelong Agents, ICLR 2026 Workshop World Models, COLM 2026 Workshop on Lifelong Agents, COLM 2026 Workshop Efficient Reasoning, NeurIPS 2026 Workshop on Physical Understanding for Decision-Making.

πŸ“– Education

πŸ’Ό Experiences