About Me
- I am a Senior Researcher at Microsoft Research Asia (MSRA). I received my Ph.D. from Department of Electronic and Computer Engineering (ECE) at Hong Kong University of Science and Technology (HKUST), supervised by Prof. Jun Zhang. I received my B.Eng in School of Electronic Information and Enginnering from South China University of Science and Technology (SCUT) in 2021.
📌 My research interests lie in Text Understanding, Multimodal Understanding, and Multimodal Generation.
✉️ Welcome to contact me for any discussion and cooperation!
💥 💥 I am looking for self-motivated interns at MSRA. If you are interested in the pre-training or post-training of large models (including LLMs, VLMs, image/video generation, and world models), please feel free to contact me at xinjiezhang@microsoft.com (cc xzhangga@connect.ust.hk). 💥 💥
🔥 News
- [2026/07] Our paper “Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model” was released. [Paper][Project Page][Code][Models]
- [2026/07] Our paper “Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing” was released. [Paper][Project Page][Code][Models]
- [2026/07] Our paper “SciForma: Structure-Faithful Generation of Scientific Diagrams” was accepted to SIGGRAPH Asia 2026. [Paper][Code]
- [2026/02] Our paper “MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model” was accepted to CVPR 2026.
- [2026/01] Our paper “SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation” was accepted to ICLR 2026. [Paper][Code]
- [2026/01] Our paper “Dynamics-Aware Gaussian Splatting Streaming Towards Fast On-the-Fly 4D Reconstruction” was accepted to IEEE Transactions on Visualization and Computer Graphics (TVCG). [Paper][Project Page]
- [2025/12] Our paper “Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling” was released. [Paper]
- [2025/11] Our paper “GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting” was accepted to AAAI 2026. [Paper][Code]
- [2025/06] We release our latest unified multimodal understanding and generation foundation model Ovis-U1. Have a try! [Paper][Project Page] [Demo]
- [2025/06] Our paper “MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes” was accepted to ICCV 2025. [Paper][Code]
- [2025/05] Our paper “Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities” was released. [Paper][Project Page]
- [2025/05] Our paper “HarmoniCa: Harmonizing Training and Inference for Better Feature Cache in Diffusion Transformer Acceleration” was accepted to ICML 2025. [Paper]
- [2025/01] Our paper “Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior” was accepted to ICLR 2025. [Paper][Code]
- [2025/01] Our paper “PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video Recognition” was accepted to ICLR 2025. [Paper]
💼 Work Experience
Senior Researcher @ Microsoft Research Asia (MSRA) 🇭🇰 Hong Kong, Jul. 2025 - Present - Manager: Dr. Yan Lu
🏢 Internship Experience
Research Intern @ AI Businesss, Alibaba 🇨🇳 Hangzhou, Dec. 2024 - May 2025 - Mentor: Dr. Guohua Wang, Mr. Qingguo Chen
Research Intern @ 2050 Research, Skywork AI 🇸🇬 Singapore, Jun. 2024 - Dec. 2024 - Mentor: Dr. Yifan Zhang, Advisor: Prof. Shuicheng Yan
Research Intern @ Model Toolchain Team, SenseTime Research 🇨🇳 Beijing, Jan. 2024 - May 2024 - Mentor: Dr. Ruihao Gong
Research Intern @ ISP and Codec Team, SenseTime Research 🇨🇳 Beijing, May 2023 - Jan. 2024 - Mentor: Dr. Ren Yang
📝 Selected Publications
Refer to my Google Scholar Profile for full publication list.
(* Equal contribution; † Corresponding author; ‡ Project lead.)
- Multimodal Understanding and Generation:
- S. Yang*, K. Zhang*, Z. Jia*, J. Guo*, Y. Shen*‡, X. Zhang*‡, X. Zhang, H. Wang, X. Li, P. Zhang, X. An, Y. Xie, Z. Liu, X. Guo, J. Li, S. Zheng, J. Wang, Z. Guo, W. Xie, Z. Zheng, Y. Luo, B. Li, and Y. Lu, “Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model”, preprint, Jul. 2026. [Paper][Project Page][Code][Models]
- X. Zhang*‡, P. Zhang*, S. Zheng*, J. Guo*, Z. Jia*, Y. Shen*, X. Guo, Y. Luo, J. Li, W. Xie, F. Pu, X. Zhang, K. Zhang, Z. Guo, T. Bi, D. Gui, Z. Liu, Z. Wen, Z. Zheng, S. Yang, X. Li, J. Wang, B. Li, and Y. Lu, “Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing”, preprint, Jul. 2026. [Paper][Project Page][Code][Models]
- Y. Luo, P. Zhang, X. Zhang†, X. Guo, Z. Lian†, and Y. Lu, “SciForma: Structure-Faithful Generation of Scientific Diagrams”, preprint, Jul. 2026. [Paper][Code]
- G. Wang, S. Zhao, X. Zhang, L. Cao, P. Zhan, L. Duan, S. Lu, M. Fu, X. Chen, J. Zhao, Y. Li, Q. Chen, “Ovis-U1 Technical Report”, preprint, Jun. 2025. [Paper][Project Page][Demo]
- X. Zhang*, J. Guo*, S. Zhao*, M. Fu, L. Duan, G. Wang, Q. Chen, Z. Xu, W. Luo, K. Zhang, “Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities”, preprint, May 2025. [Paper][Project Page]
- X. Ge, X. Zhang, T. Xu, Y. Zhang, X. Zhang, Y. Wang, and J. Zhang, “SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation,” International Conference on Learning Representations (ICLR), 2026. [Paper][Code]
- Y. Huang, Z. Wang, R. Gong, J. Liu, X. Zhang, J. Guo, X. Liu, and J. Zhang, “HarmoniCa: Harmonizing training and inference for better feature cache in diffusion transformer acceleration,” International Conference on Machine Learning (ICML), Vancouver, Canada, July 2025. [Paper][Code]
- Neural Data Representation:
- Z. Liu, Y. Hu, X. Zhang, R. Song, J. Shao, Z. Lin, and J. Zhang, “Dynamics-Aware Gaussian Splatting Streaming Towards Fast On-the-Fly 4D Reconstruction,” IEEE Transactions on Visualization and Computer Graphics (TVCG), 2026. [Paper][Project Page]
- X. Zhang, Z. Liu, Y. Zhang, X. Ge, D. He, T. Xu, Y. Wang, S. Yan and J. Zhang, “MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes”, International Conference on Computer Vision (ICCV), Honolulu, Hawai’i, USA, Oct. 2025. [Paper][Code]
- T. Li, X. Zhang, X. Ge, T. Xu, D. He, J. Zhang, and Y. Wang, “GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting,” Pro. AAAI Conference on Artificial Intelligence, 2026. [Paper][Code]
- X. Zhang*, X. Ge*, T. Xu, D. He, Y. Wang, H. Qin, G. Lu, J. Geng, and J. Zhang, “GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting,” European Conference on Computer Vision (ECCV), Milano, Italy, Sept.-Oct. 2024. [Paper][Code] (* equal contribution)
- X. Zhang, R. Yang, D. He, X. Ge, T. Xu, Y. Wang, H. Qin, and J. Zhang, “Boosting Neural Representations for Videos with a Conditional Decoder,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, June 2024. [Paper][Code] (Highlight)
- Neural Data Compression:
- S. Qin, X. Zhang, Z. Liu, J. Wang, B. Chen, J. Li, Y. Ren, S.-T. Xia, and J. Zhang, “MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.
- Z. Liu, R. Song, Y. Huang, Y. Hu, X. Zhang, J. Shao, Z. Lin, and J. Zhang, “Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling,” preprint, Dec. 2025. [Paper]
- X. Zhang, S. Gao, Z. Liu, J. Shao, X. Ge, D. He, T. Xu, Y. Wang, and J. Zhang, “CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression”, Pro. AAAI Conference on Artificial Intelligence, Philadelphia, USA, Feb.-Mar. 2025. [Paper][Code]
- Z. Liu, X. Zhang, J. Shao, Z. Lin, J. Zhang, “Bidirectional Stereo Image Compression with Cross-Dimensional Entropy Model,” European Conference on Computer Vision (ECCV), Milano, Italy, Sept.-Oct. 2024. [Paper][Code]
- X. Ge, J. Luo, X. Zhang, T. Xu, G. Lu, D. He, J. Geng, Y. Wang, J. Zhang, and H. Qin, “Task-aware Encoder Control for Deep Video Compression,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, Jun. 2024. [Paper]
- X. Zhang, J. Shao, and J. Zhang, “Low-complexity Deep Video Compression with A Distributed Coding Architecture,” IEEE International Conference on Multimedia and Expo (ICME), Brisbane, Australia, Jul. 2023. [Paper][Code]
- X. Zhang, J. Shao, and J. Zhang, “LDMIC: Learning-based distributed multi-view image coding,” International Conference on Learning Representations (ICLR), Kigali, Rwanda, May 2023. [Paper][Code]
- Edge AI:
- Y. Shen, J. Shao, X. Zhang, Z. Lin, H. Pan, D. Li, J. Zhang, K. B. Letaief, “Large language models empowered autonomous edge AI for connected intelligence,” IEEE Commun. Mag., to appear. [Paper]
- J. Shao, X. Zhang, and J. Zhang, “Task-oriented communication for edge video analytics,” IEEE Trans. Wireless Commun., to appear. [Paper]
- X. Zhang, X. Zhang and W. Yang, “Joint Offloading and Resource Allocation Using Deep Reinforcement Learning in Mobile Edge Computing,” IEEE Transactions on Network Science and Engineering, vol. 9, no. 5, pp. 3454-3466, 1 Sept.-Oct. 2022. [Paper]
- X. Zhang, J. Shao, Y. Mao, and J. Zhang, “Communication-Computation Efficient Device-Edge Co-Inference via AutoML,” IEEE Global Communications Conference (GLOBECOM), Madrid, Spain, December 2021. [Paper]
🎖 Awards and Honors
- Top Reviewer, NeurIPS 2024.
- Hong Kong Postgraduate Scholarship, 2021-2025.
- Cai Jianzhong First Prize Scholarship, 2019-2020.
- Undergraduate National Scholarship, 2018-2019.
- Undergraduate National Scholarship, 2017-2018.
