About Me

I am an assistant researcher at Tsinghua University, working in the Natural Language Processing Lab (THUNLP) on multimodal large language models. I received my Ph.D. degree in June 2024 from the University of Chinese Academy of Sciences, where I was a member of LAMP, advised by Prof. Qixiang Ye. I received my B.E. degree from Wuhan University in June 2019.

My research interests are multimodal large models and multimodal data representation learning, specifically high-resolution image visual perception, multimodal video understanding, and remote sensing multimodal large models. Recently I am also interested in generative visual perception for everything in the world, as well as spatial intelligence. Welcome for discussion and collaboration, feel free to drop me an email.

I have published more than 30 papers at top-tier artificial intelligence venues, including CVPR, NeurIPS, ACL, ICCV, ECCV, IJCV, IEEE TPAMI and IEEE TNNLS, and my work has received 3,000 citations on Google Scholar. One of them was selected as a Best Paper Candidate at CVPR 2026. I serve as a reviewer for CVPR, NeurIPS, ECCV, IEEE TPAMI and IEEE TIP, where I received the Outstanding Reviewer Award at CVPR 2026.

Selected Publications

Showing all 11 selected publications.

  1. Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation

    ECCV 2026CCF-BCorresponding author

    Yichen Zhang, Da Peng, Zonghao Guo, Zijian Zhang, Xuesong Yang, Tong Sun, Shichu Sun, Yidan Zhang, Yanghao Li, Haiyan Zhao, et al.

    A unified multimodal model that separates patch detail from semantics, matching Tar-1.5B on GenEval and MMBench at 20% of the training cost.

    Citations 1

  2. LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs

    ECCV 2026CCF-BCorresponding author

    Shichu Sun, Yichen Zhang, Haolin Song, Zonghao Guo, Chi Chen, Yidan Zhang, Yuan Yao, Zhiyuan Liu, Maosong Sun

    Progressive visual compression reconfigures a pretrained ViT into ViT-UHD, matching Qwen2-VL while reducing time-to-first-token by 1.9×.

    Citations 3

  3. RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

    ACL 2026CCF-ACorresponding author

    Bingxian Wu, Yu Zhang, Zonghao Guo, Tang Liu, Chen Qian, Yuxiang Lu, Xingbo Du, Yanghao Li, Yidan Zhang, Chi Chen, Ling Yao, Chenghu Zhou, Maosong Sun

    Bootstraps remote sensing agents with distilled domain knowledge and failure-aware experience, gaining 6% accuracy on EarthBench with DeepSeek-V3.2.

  4. FlexiVideo: Variation-Aware Temporal Dynamics Modeling for Efficient Video Understanding

    CVPR 2026CCF-ACorresponding author

    Da Peng, Xuesong Yang, Zonghao Guo, Yichen Zhang, Chi Chen, Yidan Zhang, Yuan Yao, Fang Wan, Wei Ke, Maosong Sun

    Groups frames into scene segments by visual variation, cutting visual tokens by 43.5% on MotionBench while outperforming Qwen2.5-VL-3B.

  5. GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

    CVPR 2026, OralCCF-ABest Paper CandidateCorresponding author

    Peirong Zhang, Yidan Zhang, Luxiao Xu, Jinliang Lin, Zonghao Guo, Fengxiang Wang, Xue Yang, Kaiwen Wei, Lei Wang

    Reformulates remote sensing visual grounding as reward-guided tree search, locating tiny targets inside kilometre-scale scenes before conditional grounding.

    Citations 3

  6. LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic Pyramid

    AAAI 2026CCF-ACorresponding author

    Yipeng Zhang, Yifan Liu, Zonghao Guo, Yidan Zhang, Xuesong Yang, Chi Chen, Jun Song, Bo Zheng, Yuan Yao, Zhiyuan Liu, Tat-Seng Chua, Maosong Sun

    A hierarchical window transformer that builds an inverse semantic pyramid, progressively injecting low-level visual detail into high-level semantics for fine-grained perception.

    Citations 31

  7. Video-R1: Reinforcing Video Reasoning in MLLMs

    NeurIPS 2025CCF-AMost Influential Paper Top 10Corresponding author

    Kaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo, Yibing Wang, Tianshuo Peng, Junfei Wu, Xiaoying Zhang, Benyou Wang, Xiangyu Yue

    The first systematic study of the R1 paradigm for video reasoning; T-GRPO adds temporal modelling, and Video-R1-7B surpasses GPT-4o on VSI-Bench.

    Citations 469

  8. Discriminatively Matched Part Tokens for Pointly Supervised Instance Segmentation

    IJCV 2025CCF-A · IF 10.3 · CAS Q1First author

    Zonghao Guo, Fang Wan, Mingxiang Liao, Yidan Zhang, Qixiang Ye

    Allocates a token per object part and re-estimates it with deformable part classifiers, improving point-supervised segmentation by 2.0% mAP50 on PASCAL VOC.

    Citations 3

  9. LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

    ECCV 2024CCF-BFirst author

    Zonghao Guo, Ruyi Xu, Yuan Yao, Junbo Cui, Zanlin Ni, Chunjiang Ge, Tat-Seng Chua, Zhiyuan Liu, Maosong Sun, Gao Huang

    Perceives images in any aspect ratio and native high resolution through adaptive slicing, rather than squeezing them into a fixed low-resolution square.

    Citations 285

  10. Conformer: Local Features Coupling Global Representations for Recognition and Detection

    IEEE TPAMI 2023CCF-A · IF 20.4 · CAS Q1 TopSecond author

    Zhiliang Peng, Zonghao Guo, Wei Huang, Yaowei Wang, Lingxi Xie, Jianbin Jiao, Qixiang Ye

    Couples local convolutional features with global self-attention representations in a dual-branch backbone for recognition and detection.

    Citations 1,360

  11. Convex-Hull Feature Adaptation for Oriented and Densely Packed Object Detection

    IEEE TCSVT 2022CCF-B · IF 10.8 · CAS Q1 TopFirst author

    Zonghao Guo, Xiaosong Zhang, Chang Liu, Xiangyang Ji, Jianbin Jiao, Qixiang Ye

    Represents oriented and densely packed objects with convex hulls instead of axis-aligned boxes; the conference version appeared at CVPR 2021.

    Citations 373

Honors and Awards

  • Best Paper Candidate, CVPR 2026 (GeoViS).
  • Outstanding Reviewer, CVPR 2026.
  • Excellent Student Scholarship, Chinese Academy of Sciences, 2020.

Experience and Education

  1. 2024 — Present

    Assistant Researcher, Tsinghua University

    Natural Language Processing Lab (THUNLP), working on multimodal large language models. Beijing, China.

  2. 2019 — Jun 2024

    Ph.D., University of Chinese Academy of Sciences

    LAMP. Advisor: Prof. Qixiang Ye. Beijing, China.

  3. Jun 2019

    B.E., Wuhan University

    Bachelor of Engineering. Wuhan, China.

Contact

For collaborations or enquiries, please write to guozonghao96@outlook.com.