About Me

📌 My research interests lie in Text Understanding, Multimodal Understanding, and Multimodal Generation.

✉️ Welcome to contact me for any discussion and cooperation!

💥 💥 I am looking for self-motivated interns at MSRA. If you are interested in the pre-training or post-training of large models (including LLMs, VLMs, image/video generation, and world models), please feel free to contact me at xinjiezhang@microsoft.com (cc xzhangga@connect.ust.hk). 💥 💥

🔥 News

  • [2026/07] Our paper “Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model” was released. [Paper][Project Page][Code][Models]
  • [2026/07] Our paper “Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing” was released. [Paper][Project Page][Code][Models]
  • [2026/07] Our paper “SciForma: Structure-Faithful Generation of Scientific Diagrams” was accepted to SIGGRAPH Asia 2026. [Paper][Code]
  • [2026/02] Our paper “MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model” was accepted to CVPR 2026.
  • [2026/01] Our paper “SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation” was accepted to ICLR 2026. [Paper][Code]
  • [2026/01] Our paper “Dynamics-Aware Gaussian Splatting Streaming Towards Fast On-the-Fly 4D Reconstruction” was accepted to IEEE Transactions on Visualization and Computer Graphics (TVCG). [Paper][Project Page]
  • [2025/12] Our paper “Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling” was released. [Paper]
  • [2025/11] Our paper “GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting” was accepted to AAAI 2026. [Paper][Code]
  • [2025/06] We release our latest unified multimodal understanding and generation foundation model Ovis-U1. Have a try! [Paper][Project Page] [Demo]
  • [2025/06] Our paper “MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes” was accepted to ICCV 2025. [Paper][Code]
  • [2025/05] Our paper “Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities” was released. [Paper][Project Page]
  • [2025/05] Our paper “HarmoniCa: Harmonizing Training and Inference for Better Feature Cache in Diffusion Transformer Acceleration” was accepted to ICML 2025. [Paper]
  • [2025/01] Our paper “Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior” was accepted to ICLR 2025. [Paper][Code]
  • [2025/01] Our paper “PvNeXt: Rethinking Network Design and Temporal Motion for Point Cloud Video Recognition” was accepted to ICLR 2025. [Paper]

💼 Work Experience

🏢 Internship Experience

  • Research Intern@ AI Businesss, Alibaba🇨🇳 Hangzhou, Dec. 2024 - May 2025
  • Research Intern@ 2050 Research, Skywork AI🇸🇬 Singapore, Jun. 2024 - Dec. 2024
  • Research Intern@ Model Toolchain Team, SenseTime Research🇨🇳 Beijing, Jan. 2024 - May 2024
  • Research Intern@ ISP and Codec Team, SenseTime Research🇨🇳 Beijing, May 2023 - Jan. 2024

📝 Selected Publications

Refer to my Google Scholar Profile for full publication list.

(* Equal contribution; † Corresponding author; ‡ Project lead.)

  • Multimodal Understanding and Generation:
    • S. Yang*, K. Zhang*, Z. Jia*, J. Guo*, Y. Shen*‡, X. Zhang*‡, X. Zhang, H. Wang, X. Li, P. Zhang, X. An, Y. Xie, Z. Liu, X. Guo, J. Li, S. Zheng, J. Wang, Z. Guo, W. Xie, Z. Zheng, Y. Luo, B. Li, and Y. Lu, “Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model”, preprint, Jul. 2026. [Paper][Project Page][Code][Models]
    • X. Zhang*‡, P. Zhang*, S. Zheng*, J. Guo*, Z. Jia*, Y. Shen*, X. Guo, Y. Luo, J. Li, W. Xie, F. Pu, X. Zhang, K. Zhang, Z. Guo, T. Bi, D. Gui, Z. Liu, Z. Wen, Z. Zheng, S. Yang, X. Li, J. Wang, B. Li, and Y. Lu, “Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing”, preprint, Jul. 2026. [Paper][Project Page][Code][Models]
    • Y. Luo, P. Zhang, X. Zhang†, X. Guo, Z. Lian†, and Y. Lu, “SciForma: Structure-Faithful Generation of Scientific Diagrams”, preprint, Jul. 2026. [Paper][Code]
    • G. Wang, S. Zhao, X. Zhang, L. Cao, P. Zhan, L. Duan, S. Lu, M. Fu, X. Chen, J. Zhao, Y. Li, Q. Chen, “Ovis-U1 Technical Report”, preprint, Jun. 2025. [Paper][Project Page][Demo]
    • X. Zhang*, J. Guo*, S. Zhao*, M. Fu, L. Duan, G. Wang, Q. Chen, Z. Xu, W. Luo, K. Zhang, “Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities”, preprint, May 2025. [Paper][Project Page]
    • X. Ge, X. Zhang, T. Xu, Y. Zhang, X. Zhang, Y. Wang, and J. Zhang, “SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation,” International Conference on Learning Representations (ICLR), 2026. [Paper][Code]
    • Y. Huang, Z. Wang, R. Gong, J. Liu, X. Zhang, J. Guo, X. Liu, and J. Zhang, “HarmoniCa: Harmonizing training and inference for better feature cache in diffusion transformer acceleration,” International Conference on Machine Learning (ICML), Vancouver, Canada, July 2025. [Paper][Code]
  • Neural Data Representation:
    • Z. Liu, Y. Hu, X. Zhang, R. Song, J. Shao, Z. Lin, and J. Zhang, “Dynamics-Aware Gaussian Splatting Streaming Towards Fast On-the-Fly 4D Reconstruction,” IEEE Transactions on Visualization and Computer Graphics (TVCG), 2026. [Paper][Project Page]
    • X. Zhang, Z. Liu, Y. Zhang, X. Ge, D. He, T. Xu, Y. Wang, S. Yan and J. Zhang, “MEGA: Memory-Efficient 4D Gaussian Splatting for Dynamic Scenes”, International Conference on Computer Vision (ICCV), Honolulu, Hawai’i, USA, Oct. 2025. [Paper][Code]
    • T. Li, X. Zhang, X. Ge, T. Xu, D. He, J. Zhang, and Y. Wang, “GaussianImage++: Boosted Image Representation and Compression with 2D Gaussian Splatting,” Pro. AAAI Conference on Artificial Intelligence, 2026. [Paper][Code]
    • X. Zhang*, X. Ge*, T. Xu, D. He, Y. Wang, H. Qin, G. Lu, J. Geng, and J. Zhang, “GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting,” European Conference on Computer Vision (ECCV), Milano, Italy, Sept.-Oct. 2024. [Paper][Code] (* equal contribution)
    • X. Zhang, R. Yang, D. He, X. Ge, T. Xu, Y. Wang, H. Qin, and J. Zhang, “Boosting Neural Representations for Videos with a Conditional Decoder,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, June 2024. [Paper][Code] (Highlight)
  • Neural Data Compression:
    • S. Qin, X. Zhang, Z. Liu, J. Wang, B. Chen, J. Li, Y. Ren, S.-T. Xia, and J. Zhang, “MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.
    • Z. Liu, R. Song, Y. Huang, Y. Hu, X. Zhang, J. Shao, Z. Lin, and J. Zhang, “Feed-Forward 3D Gaussian Splatting Compression with Long-Context Modeling,” preprint, Dec. 2025. [Paper]
    • X. Zhang, S. Gao, Z. Liu, J. Shao, X. Ge, D. He, T. Xu, Y. Wang, and J. Zhang, “CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression”, Pro. AAAI Conference on Artificial Intelligence, Philadelphia, USA, Feb.-Mar. 2025. [Paper][Code]
    • Z. Liu, X. Zhang, J. Shao, Z. Lin, J. Zhang, “Bidirectional Stereo Image Compression with Cross-Dimensional Entropy Model,” European Conference on Computer Vision (ECCV), Milano, Italy, Sept.-Oct. 2024. [Paper][Code]
    • X. Ge, J. Luo, X. Zhang, T. Xu, G. Lu, D. He, J. Geng, Y. Wang, J. Zhang, and H. Qin, “Task-aware Encoder Control for Deep Video Compression,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, USA, Jun. 2024. [Paper]
    • X. Zhang, J. Shao, and J. Zhang, “Low-complexity Deep Video Compression with A Distributed Coding Architecture,” IEEE International Conference on Multimedia and Expo (ICME), Brisbane, Australia, Jul. 2023. [Paper][Code]
    • X. Zhang, J. Shao, and J. Zhang, “LDMIC: Learning-based distributed multi-view image coding,” International Conference on Learning Representations (ICLR), Kigali, Rwanda, May 2023. [Paper][Code]
  • Edge AI:
    • Y. Shen, J. Shao, X. Zhang, Z. Lin, H. Pan, D. Li, J. Zhang, K. B. Letaief, “Large language models empowered autonomous edge AI for connected intelligence,” IEEE Commun. Mag., to appear. [Paper]
    • J. Shao, X. Zhang, and J. Zhang, “Task-oriented communication for edge video analytics,” IEEE Trans. Wireless Commun., to appear. [Paper]
    • X. Zhang, X. Zhang and W. Yang, “Joint Offloading and Resource Allocation Using Deep Reinforcement Learning in Mobile Edge Computing,” IEEE Transactions on Network Science and Engineering, vol. 9, no. 5, pp. 3454-3466, 1 Sept.-Oct. 2022. [Paper]
    • X. Zhang, J. Shao, Y. Mao, and J. Zhang, “Communication-Computation Efficient Device-Edge Co-Inference via AutoML,” IEEE Global Communications Conference (GLOBECOM), Madrid, Spain, December 2021. [Paper]

🎖 Awards and Honors

  • Top Reviewer, NeurIPS 2024.
  • Hong Kong Postgraduate Scholarship, 2021-2025.
  • Cai Jianzhong First Prize Scholarship, 2019-2020.
  • Undergraduate National Scholarship, 2018-2019.
  • Undergraduate National Scholarship, 2017-2018.