📄 About Me
I am a master’s student at the Institute of Automation, Chinese Academy of Sciences (CASIA), affiliated with the National Laboratory of Pattern Recognition, supervised by Prof. Haiyun Guo and under the leadership of Prof. Jinqiao Wang. I will join Shanghai Jiao Tong University as an incoming Ph.D. student in September 2026 through a joint-training program with Shanghai Innovation Institute, supervised by Prof. Junchi Yan.
My research focuses on two directions:
- Unified Multimodal Models (UMM): unified multimodal models that use generation to assist perception, reasoning, and understanding, especially for spatial reasoning, degraded image understanding, dynamic-resolution reasoning, and spatially grounded text-to-image evaluation.
- Universal Multimodal Embedding (UME): universal multimodal embedding and retrieval, reasoning-enhanced embedding, fine-grained instance retrieval, composed image retrieval, and fine-grained visual classification.
📄 关于我
我目前是中国科学院自动化研究所(CASIA)硕士研究生,隶属于模式识别国家重点实验室,导师为郭海云教授,并在王金桥教授指导下开展科研工作。我将于 2026 年 9 月加入上海交通大学攻读博士学位,与上海创智学院联培,导师为严骏驰教授。
我的研究主要围绕两个方向:
- Unified Multimodal Models (UMM):研究利用 generation 辅助 perception、reasoning 和 understanding 的 unified multimodal models,重点关注 spatial reasoning、degraded image understanding、dynamic-resolution reasoning 和 spatially grounded text-to-image evaluation。
- Universal Multimodal Embedding (UME):研究 universal multimodal embedding and retrieval、reasoning-enhanced embedding、fine-grained instance retrieval、composed image retrieval 和 fine-grained visual classification。
🔥 News
- 2026.09: I will join Shanghai Jiao Tong University as a Ph.D. student through a joint-training program with Shanghai Innovation Institute.
- 2026.07: 🎉🎉 Our paper “PLUME” was accepted to ACM MM 2026!
- 2026: 🎉🎉 Our paper “COOPER” received the CVPR 2026 Compute Transparency Champion award!
- 2026.02: 🎉🎉 Our papers “COOPER”, “WISER”, and “ReCALL” were accepted to CVPR 2026!
- 2025.07: Started research internship at Baidu Wenxin (ERNIE Bot) team.
- 2025.07: 🎉🎉 Our paper “Referring Expression Instance Retrieval and A Strong End-to-End Baseline” was accepted to ACM MM 2025!
- 2026.09: 将加入上海交通大学攻读博士学位,与上海创智学院联培。
- 2026.07: 🎉🎉 论文 “PLUME” 被 ACM MM 2026 录用!
- 2026: 🎉🎉 论文 “COOPER” 获得 CVPR 2026 Compute Transparency Champion award!
- 2026.02: 🎉🎉 论文 “COOPER”、“WISER” 和 “ReCALL” 被 CVPR 2026 录用!
- 2025.07: 开始在 Baidu Wenxin (ERNIE Bot) 团队科研实习。
- 2025.07: 🎉🎉 论文 “Referring Expression Instance Retrieval and A Strong End-to-End Baseline” 被 ACM MM 2025 录用!
💻 Internships
- 2025.07 - Present, Baidu - Wenxin (ERNIE Bot) Team, Beijing, China.
- Participated in ERNIE 5.0 multimodal pretraining, including data processing, cleaning, quality evaluation, data-mixture control, and benchmark-based model analysis.
- Worked on ERNIE-One unified understanding-generation model training. I explored dual-encoder + VAE and unified-encoder + Pixel Diffusion architectures, and participated in pretraining, mid-training, alignment, SFT, and RL stages.
- Built data and training pipelines for thinking-then-generation and interleaved image-text generation, and studied how architecture choices, scaling, and data composition affect both understanding and generation.
- 2025.01 - 2025.06, Zidongtaichu - Foundation Model Research Center, Beijing, China.
- Worked on multimodal embedding and local retrieval modules for foundation-model applications.
- 2025.07 - Present, Baidu - Wenxin (ERNIE Bot) Team, 北京,中国。
- 参与 ERNIE 5.0 multimodal pretraining,负责数据处理、数据清洗、质量评估、data-mixture control 以及基于 benchmark 的模型效果分析。
- 参与 ERNIE-One unified understanding-generation model 训练,探索 dual-encoder + VAE 与 unified-encoder + Pixel Diffusion 两类架构,并参与 pretraining、mid-training、alignment、SFT 和 RL 阶段。
- 构建 thinking-then-generation 与 interleaved image-text generation 的数据和训练 pipeline,研究 architecture choices、scaling 和 data composition 对 understanding 与 generation 能力的影响。
- 2025.01 - 2025.06, Zidongtaichu - Foundation Model Research Center, 北京,中国。
- 参与 multimodal embedding 和 local retrieval 模块研发,服务 foundation-model 应用场景。
🧠 Unified Multimodal Models (UMM)

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
Zefeng Zhang*, Xiangzhao Hao*, Hengzhu Tang, Zhenyu Zhang, Jiawei Sheng, Xiaodong Li, Zhenyang Li, Li Gao, Daiting Shi, Dawei Yin, Tingwen Liu
- We propose a cooperative perception-reasoning unified framework for spatial intelligence, where the model autonomously generates auxiliary visual information, such as depth maps and segmentation maps, as part of a multimodal chain-of-thought.
- Designed a SFT+GRPO two-stage framework with Cooperative Perception-Reasoning Reward, achieving 6.91% average improvement on spatial reasoning benchmarks.
- Received the CVPR 2026 Compute Transparency Champion award.
- 提出面向 spatial intelligence 的 cooperative perception-reasoning unified framework,使模型能够在 multimodal chain-of-thought 中按需生成 depth maps、segmentation maps 等辅助视觉信息。
- 设计 SFT+GRPO 两阶段训练框架和 Cooperative Perception-Reasoning Reward,在 spatial reasoning benchmarks 上取得 6.91% 平均提升。
- 获得 CVPR 2026 Compute Transparency Champion award。

CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
Xiangzhao Hao*, Zefeng Zhang*, Zhenyu Zhang, Linhao Yu, Yao Chen, Yiqian Zhang, Haiyun Guo, Shuohuan Wang, Yu Sun
- We explore generation-assisted understanding for degraded images, enabling unified multimodal models to decide when to restore visual evidence and when to directly answer.
- Designed Latent Representation Bridge and Interleaved GRPO to connect visual generation with answer correctness under degradation.
- 研究 degraded images 场景下的 generation-assisted understanding,使 unified multimodal models 能够判断何时恢复视觉证据、何时直接回答。
- 设计 Latent Representation Bridge 和 Interleaved GRPO,将 visual generation 与 degradation 条件下的 answer correctness 连接起来。

DynEyes: The Default Resolution Is Not the Best
Anonymous Author(s)
- We investigate sample-level resolution preference in MLLMs and show that the original image resolution is often not the best choice for multimodal reasoning.
- Proposed a reinforcement learning framework that predicts suitable resolution scales for each image-question pair and transfers across MLLMs.
- 研究 MLLMs 中 sample-level resolution preference,发现原始图像分辨率并不总是 multimodal reasoning 的最佳选择。
- 提出 reinforcement learning framework,为每个 image-question pair 预测合适的 resolution scales,并可迁移到其他 MLLMs。

Can Text-to-Image Models Draw from the Right Frame of Reference?
Zheyuan Gu, Ruihang Li, Yong Huang, Yiqian Zhang, Xiangzhao Hao, Jiaxin Niu, Jiahao Hu, Zhenyu Zhang
- We introduce FoR-T2I, a benchmark for evaluating whether text-to-image models correctly follow object-relative spatial instructions rather than defaulting to camera-view coordinates.
- Across 22 text-to-image models, frame-of-reference prompts remain substantially harder than matched camera-view prompts.
- 提出 FoR-T2I benchmark,用于评估 text-to-image models 是否能够正确遵循 object-relative spatial instructions,而不是默认使用 camera-view coordinates。
- 在 22 个 text-to-image models 上,frame-of-reference prompts 相比匹配的 camera-view prompts 明显更具挑战性。
🔎 Universal Multimodal Embedding (UME)

PLUME: Latent Reasoning Based Universal Multimodal Embedding
Chenwei He*, Xiangzhao Hao*, Tianyu Yang*, Yuxiang Ma, Yuheng Jia, Lingxiang Wu, Chaoyang Zhao, Haiyun Guo, Jinqiao Wang
- We internalize explicit CoT into short latent reasoning trajectories, reducing hundreds of generated tokens to fewer than 10 latent steps.
- PLUME achieves strong performance on MMEB-v2 while delivering over 30x faster inference than explicit-CoT retrieval baselines.
- 将 explicit CoT 内化为短步 latent reasoning trajectories,将数百个 generated tokens 压缩为少于 10 个 latent steps。
- PLUME 在 MMEB-v2 上取得较强性能,同时相比 explicit-CoT retrieval baselines 实现超过 30x 的 inference 加速。

TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
Xiangzhao Hao*, Shijie Wang*, Tianyu Yang*, Tianyue Wang, Haiyun Guo, Jinqiao Wang
- We propose a “reason-then-encode” retrieval paradigm that integrates task-adaptive CoT generation with discriminative representation learning.
- Built M-BEIR-CoT and achieved state-of-the-art performance on M-BEIR, with adaptive reasoning activation for complex queries.
- 提出 “reason-then-encode” retrieval paradigm,将 task-adaptive CoT generation 与 discriminative representation learning 结合。
- 构建 M-BEIR-CoT,在 M-BEIR 上取得 state-of-the-art performance,并能够对复杂查询自适应激活 reasoning。

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding
Shijie Wang*, Xiangzhao Hao*, Yueti Li, Guangyu Cao, Xinyu Tang, Haiyun Guo
- We expand retrieval computation along model depth instead of generating additional reasoning tokens, using parameter-shared recurrent retrieval-forming blocks.
- Learnable Retrieval Registers accumulate retrieval-specific evidence across loops, running 44.9x faster than UME-R1 and 1.5x faster than PLUME.
- 不通过生成额外 reasoning tokens,而是沿 model depth 扩展 retrieval computation,使用 parameter-shared recurrent retrieval-forming blocks。
- Learnable Retrieval Registers 在循环中累积 retrieval-specific evidence,相比 UME-R1 快 44.9x,相比 PLUME 快 1.5x。

Tianyue Wang, Leigang Qu, Tianyu Yang, Xiangzhao Hao, Yifan Xu, Haiyun Guo, Jinqiao Wang
- We propose WISER, a training-free framework for zero-shot composed image retrieval with wider search, deeper thinking, and adaptive fusion.
- Achieves 45% relative improvement on CIRCO mAP@5 and 57% on CIRR Recall@1 over prior training-free methods.
- 提出 WISER,一个用于 zero-shot composed image retrieval 的 training-free framework,包含 wider search、deeper thinking 和 adaptive fusion。
- 相比已有 training-free methods,在 CIRCO mAP@5 上取得 45% 相对提升,在 CIRR Recall@1 上取得 57% 相对提升。

ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
Tianyu Yang, Chenwei He, Xiangzhao Hao, Tianyue Wang, Jiarui Guo, Haiyun Guo, Leigang Qu, Jinqiao Wang, Tat-Seng Chua
- We diagnose capability degradation when generative MLLMs are compressed into discriminative retrievers.
- ReCALL recalibrates retrieval ability through hard negative mining, corrective instruction generation, and continual contrastive refinement.
- 诊断 generative MLLMs 被压缩为 discriminative retrievers 时出现的 capability degradation。
- ReCALL 通过 hard negative mining、corrective instruction generation 和 continual contrastive refinement 重新校准 retrieval ability。

Referring Expression Instance Retrieval and A Strong End-to-End Baseline
Xiangzhao Hao, Kuan Zhu, Hongyu Guo, Haiyun Guo, Ning Jiang, Quan Lu, Ming Tang, Jinqiao Wang
- We introduce REIR, retrieving and localizing a specific object instance from a gallery based on an instance-level natural-language description.
- Constructed REIRCOCO and proposed the end-to-end dual-stream baseline CLARE.
- 提出 REIR,根据 instance-level natural-language description 从 gallery 中检索并定位特定 object instance。
- 构建 REIRCOCO,并提出 end-to-end dual-stream baseline CLARE。

Hongyu Guo, Xiangzhao Hao*, Jiarui Guo, Haiyun Guo, Jinqiao Wang, Tat-Seng Chua
- We reformulate few-shot fine-grained visual classification as a training-free multimodal retrieval problem.
- Designed CDV-Captioner for structured attribute-aware descriptions and surpassed few-shot SOTA across 12 FGVC benchmarks.
- 将 few-shot fine-grained visual classification 重构为 training-free multimodal retrieval problem。
- 设计 CDV-Captioner 生成 structured attribute-aware descriptions,在 12 个 FGVC benchmarks 上超过 few-shot SOTA。

KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
Linhao Yu*, Tianmeng Yang*, Siyu Ding*, Renren Jin, Naibin Gu, Xiangzhao Hao, Shuaiyi Nie, Deyi Xiong, Weichong Yin, Yu Sun, Hua Wu
- We study minimal-sufficient knowledge guidance for reinforcement learning on reasoning tasks.
- KnowRL constructs compact, interaction-aware knowledge-point subsets to reduce reward sparsity without adding redundant hints.
- 研究 reasoning tasks 中用于 reinforcement learning 的 minimal-sufficient knowledge guidance。
- KnowRL 构建紧凑且 interaction-aware 的 knowledge-point subsets,以减少 reward sparsity,同时避免引入冗余 hints。

Test-Time Curriculum for Open-Set AIGC Detection
Anonymous Author(s)
- We propose a model-agnostic test-time adaptation framework for open-set AI-generated image detection under unseen generator shifts.
- TTC progressively adapts detectors from reliable pseudo-labeled samples to harder informative cases, with cross-scale pseudo-label refinement.
- 提出 model-agnostic test-time adaptation framework,用于 unseen generator shifts 下的 open-set AI-generated image detection。
- TTC 从 reliable pseudo-labeled samples 逐步过渡到更困难且信息量更高的样本,并结合 cross-scale pseudo-label refinement。
📖 Educations
- 2026.09 incoming, Ph.D. student, School of Artificial Intelligence, Shanghai Jiao Tong University; joint-training program with Shanghai Innovation Institute.
- 2023.09 - 2026.06, M.Eng. in Artificial Intelligence, Institute of Automation, University of Chinese Academy of Sciences. GPA: 3.76/4.00.
- 2019.09 - 2023.06, B.Eng. in Computer Science and Technology, School of Intelligence and Computing, Tianjin University. Ranked 6/139, GPA: 3.87/4.00.
- 2026.09 incoming, 博士研究生,上海交通大学人工智能学院;与上海创智学院联培。
- 2023.09 - 2026.06, 工程硕士,人工智能,中国科学院自动化研究所,中国科学院大学。GPA: 3.76/4.00。
- 2019.09 - 2023.06, 工学学士,计算机科学与技术,智能与计算学部,天津大学。专业排名 6/139,GPA: 3.87/4.00。
🎖 Honors and Awards
- 2026 CVPR 2026 Compute Transparency Champion award for COOPER
- 2025.10 First Place and Individual Excellence Award, Beijing Zhongguancun Academy Hackathon
- 2024.06 UCAS Outstanding Student (三好学生) and Excellent Student Leader (优秀学生干部)
- 2023.09 Second Prize in the National Final of the Loongson Cup (National Student Computer System Capability Challenge)
- 2023.07 Huawei MindSpore Scholarship
- 2023.06 Tianjin Outstanding Graduate (天津市优秀毕业生), Tianjin University Outstanding Student
- 2023.03 Meritorious Winner (一等奖), Mathematical Contest in Modeling (MCM/ICM)
- 2022.11 Silver Award, 7th China Internet+ Innovation and Entrepreneurship Competition (Tianjin)
- 2022.09 Second Prize, North China Five-Province Computer Application Competition
- 2026 COOPER 获得 CVPR 2026 Compute Transparency Champion award
- 2025.10 北京中关村学院黑客马拉松第一名、个人优胜奖
- 2024.06 中国科学院大学三好学生、优秀学生干部
- 2023.09 龙芯杯全国大学生计算机系统能力培养大赛全国总决赛二等奖
- 2023.07 Huawei MindSpore Scholarship
- 2023.06 天津市优秀毕业生、天津大学优秀学生
- 2023.03 Mathematical Contest in Modeling (MCM/ICM) Meritorious Winner
- 2022.11 第七届中国国际“互联网+”大学生创新创业大赛天津赛区银奖
- 2022.09 华北五省大学生计算机应用大赛二等奖