PH.D. STUDENT · SUN YAT-SEN UNIVERSITY

Mengzhao Wang 王盟召

I am a first-year Ph.D. student at the School of Intelligent Systems Engineering, Sun Yat-sen University (Shenzhen Campus), advised by Prof. Yanli Ji. Previously, I was a master's student at the Faculty of Information Engineering and Automation, Kunming University of Science and Technology, advised by Prof. Huafeng Li.

目前就读于中山大学深圳校区智能工程学院,研究兴趣包括高效多模态大模型推理、具身智能与视觉语言模型。

Research / 研究方向

My research focuses on efficient multimodal large-model reasoning, embodied intelligence, and vision-language models.

01

Efficient Multimodal Large-Model Reasoning

高效多模态大模型推理

02

Embodied Intelligence

具身智能

03

Vision-Language Models

视觉语言模型

News / 动态

  • Our work on replay-free visual revisiting for interleaved multimodal reasoning was released on arXiv.
  • Three papers on visual grounding and video-language understanding were published in IEEE TMM and IEEE TIP.
  • Started the Ph.D. program at Sun Yat-sen University, Shenzhen Campus.

Selected Publications / 代表论文

Model architecture for multi-event video-text retrieval and grounding

IEEE Transactions on Image Processing (TIP) · 2025

Disentangling Inter- and Intra-Video Relations for Multi-Event Video-Text Retrieval and Grounding

Mengzhao Wang, Huafeng Li, Yang Zhang, Jing Li, Dacheng Tao, Z. Yu

Disentangles inter- and intra-video relations to improve multi-event video-text retrieval and grounding.

Model architecture for joint video paragraph retrieval and grounding

IEEE Transactions on Multimedia (TMM) · 2025

Dual-Task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding

Mengzhao Wang, Huafeng Li, Yang Zhang, Jing Li, Meng Xie, Dacheng Tao

Jointly learns paragraph retrieval and grounding through mutually reinforcing dual-task training.

Model architecture for phrase decoupling visual grounding

IEEE Transactions on Multimedia (TMM) · 2025

Phrase Decoupling Cross-Modal Hierarchical Matching and Progressive Position Correction for Visual Grounding

Meng Xie, Mengzhao Wang, Huafeng Li, Yang Zhang, Dacheng Tao, Z. Yu

Combines phrase-decoupled cross-modal hierarchical matching with progressive position correction for visual grounding.

Education / 教育经历

2025–Present

Sun Yat-sen University, Shenzhen Campus

Ph.D. student, School of Intelligent Systems Engineering · Advisor: Prof. Yanli Ji

2022–2025

Kunming University of Science and Technology

Master's student, Faculty of Information Engineering and Automation · Advisor: Prof. Huafeng Li

Visit Stats · Total Visits -- 访问统计 · 总访问量 --