PH.D. STUDENT · SUN YAT-SEN UNIVERSITY
Mengzhao Wang 王盟召
I am a first-year Ph.D. student at the School of Intelligent Systems Engineering, Sun Yat-sen University (Shenzhen Campus), advised by Prof. Yanli Ji. Previously, I was a master's student at the Faculty of Information Engineering and Automation, Kunming University of Science and Technology, advised by Prof. Huafeng Li.
目前就读于中山大学深圳校区智能工程学院,研究兴趣包括高效多模态大模型推理、具身智能与视觉语言模型。
Research / 研究方向
My research focuses on efficient multimodal large-model reasoning, embodied intelligence, and vision-language models.
01
Efficient Multimodal Large-Model Reasoning
高效多模态大模型推理
02
Embodied Intelligence
具身智能
03
Vision-Language Models
视觉语言模型
News / 动态
- Our work on replay-free visual revisiting for interleaved multimodal reasoning was released on arXiv.
- Three papers on visual grounding and video-language understanding were published in IEEE TMM and IEEE TIP.
- Started the Ph.D. program at Sun Yat-sen University, Shenzhen Campus.
Selected Publications / 代表论文
arXiv preprint · 2026
Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning
Mengzhao Wang, Yanli Ji, Wangmeng Zuo, Peng Ye, Chongjun Tu
A replay-free cache-reuse framework that revisits the right visual evidence during interleaved multimodal reasoning.
IEEE Transactions on Image Processing (TIP) · 2025
Disentangling Inter- and Intra-Video Relations for Multi-Event Video-Text Retrieval and Grounding
Mengzhao Wang, Huafeng Li, Yang Zhang, Jing Li, Dacheng Tao, Z. Yu
Disentangles inter- and intra-video relations to improve multi-event video-text retrieval and grounding.
IEEE Transactions on Multimedia (TMM) · 2025
Dual-Task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
Mengzhao Wang, Huafeng Li, Yang Zhang, Jing Li, Meng Xie, Dacheng Tao
Jointly learns paragraph retrieval and grounding through mutually reinforcing dual-task training.
IEEE Transactions on Multimedia (TMM) · 2025
Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual Grounding
Meng Xie, Mengzhao Wang, Huafeng Li, Yang Zhang, Dacheng Tao, Z. Yu
Combines phrase-decoupled cross-modal hierarchical matching with progressive position correction for visual grounding.
Education / 教育经历
2025–Present
Sun Yat-sen University, Shenzhen Campus
Ph.D. student, School of Intelligent Systems Engineering · Advisor: Prof. Yanli Ji
2022–2025
Kunming University of Science and Technology
Master's student, Faculty of Information Engineering and Automation · Advisor: Prof. Huafeng Li



