👨 About Me
I received my M.Eng. degree (2026) and B.Eng. degree (2023) from MC2 Lab at Beihang University, supervised by Prof. Shengxi Li and Prof. Mai Xu. I am planning to commence my PhD studies at the Hong Kong University of Science and Technology (HKUST) in Fall 2026, under the joint supervision of Prof. Pan Wang and Prof. Yike Guo. I am broadly interested in applied computer vision, generative AI and world models.
My research interests lie in computer vision and generative AI, with a particular focus on:
- World Model
- Generative Image/Video Compression
- Image/Video Coding for Machines
💻 Experience
- 12/2025 ~ Present, Research Intern at Multimedia Lab, Taobao, Alibaba Group, supervised by Dr. Ying Chen.
- 05/2025 ~ 12/2025, Research Assistant at Tsinghua University, supervised by Dr. Tongda Xu and Prof. Yan Wang.
- 09/2023 ~ 01/2026, Earned Master’s degree in School of Electronics and Information Engineering at Beihang University, supervised by Prof. Shengxi Li and Prof. Mai Xu.
- 09/2019 ~ 06/2023, Earned Bachelor’s degree in School of Electronic Information Engineering at Beihang University.
🏆 Honors and Awards
- China National Scholarship, 2025.
- BYD Inc. Scholarship, 2025.
- Outstanding Graduate of Beijing, 2025.
- Beihang Outstanding Master Dissertation Award, 2025.
- Multiple Merit Scholarships, Beihang University, 2019-2025.
📚 Publications

Benchmarking and Enhancing VLM for Compressed Image Understanding
Zifu Zhang, Tongda Xu†, Siqi Li, Shengxi Li†, Yue Zhang, Mai Xu, Yan Wang†
Paper Code
We present a comprehensive benchmark designed to evaluate the ability of VLMs to understand and process compressed images. Based on this, we propose a lightweight VLM adaptor that enhances VLM performance on compressed images across diverse codecs and bitrate levels.

Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework
Zifu Zhang, Shengxi Li†, Xiancheng Sun, Mai Xu, Zhengyuan Liu, Jingyuan Xia
Paper Code
We propose the first successful attempt by a novel collaborative compression method based on the machine-vision-oriented compression, named as diffusion-prior based feature compression for human and machine visions (Diff-FCHM).

Hierarchical Semantic Compression for Consistent Image Semantic Restoration
Shengxi Li (Supervisor), Zifu Zhang, Mai Xu, Lai Jiang, Yufan Liu, Ce Zhu
Paper Code
We propose a novel hierarchical semantic compression (HSC) framework that purely operates within intrinsic semantic spaces from generative models, which is able to achieve efficient compression for consistent semantic restoration.

FC-FORMER: Efficient Feature Coding for Machines via a Hybrid CNN-Transformer Architecture
Zhengyuan Liu*, Zifu Zhang*, Shengxi Li†, Tao Xu, Mai Xu, Xin Deng
Paper
We propose in this paper a novel feature compression architecture based on a hybrid of Transformer and convolutional neural network (CNN) architectures, thus named as FC-Former that enhances the capability of feature extraction, achieving the state-of-the-art performances.

Continuous Patch Stitching for Block-wise Image Compression
Zifu Zhang, Shengxi Li†, Henan Liu, Mai Xu, Ce Zhu
Paper Code
We propose a novel continuous patch stitching (CPS) framework for block-wise image compression that is able to achieve seamlessly patch stitching and mathematically eliminate block artefact, thus capable of significantly reducing the required computing resources when compressing images.

Hybrid Single Input and Multiple Output Method For Compressing Features Towards Machine Vision Tasks
Zifu Zhang, Shengxi Li, Tie Liu, Mai Xu, Tao Xu, Zhenyu Guan
Paper Code
This paper introduces a simple yet effective architecture called hybrid single input and multiple output (H-SIMO) for VCM, which can significantly reduce the redundancy across scales of features.
