ICML 2026Unified multimodal learning

Generation Enhances Understanding
in Unified Multimodal Models via
Multi-Representation Generation

Pixel, depth, and segmentation.
Generation as a path to richer visual understanding.

Zihan Su1 2†* Hongyang Wei1* Kangrui Cen3* Yong Wang2‡ Guanhua Chen4
Chun Yuan1 Xiangxiang Chu2

1 Tsinghua University 2 AMAP, Alibaba Group
3 Shanghai Jiao Tong University 4 Southern University of Science and Technology
† Work done during internship at AMAP, Alibaba Group * Equal contribution ‡ Project lead
Pixel

Appearance

Image reconstruction
Depth

Geometry

Spatial relations
Segmentation

Structure

Region partitions
Generation → understanding01 / The idea

Pixel. Depth. Segmentation.

Multiple representations. A deeper understanding.

Generation as an auxiliary task for visual understanding.
Generation as an auxiliary task for visual understanding.
The idea

Generation helps models
understand what they see.

Unified multimodal models can both understand and generate images. Yet using generation to improve understanding remains less explored.

UniMRG trains models to generate pixels, depth maps, and segmentation maps alongside visual understanding tasks. These complementary signals improve fine-grained perception, reduce hallucinations, and strengthen spatial understanding.

Read the full abstract ↗
02

One model, complementary representations

Overview of UniMRG. 
      The input image is fed into the visual understanding encoder, and the UMM is jointly trained on four tasks:
      (1) Image reconstruction: reconstructin
Overview of UniMRG. The input image is fed into the visual understanding encoder, and the UMM is jointly trained on four tasks:
(1) Image reconstruction: reconstructing the input image to enhance generation capabilities.
(2) Image-to-depth: generating depth maps to learn geometric cues and spatial relations.
(3) Image-to-segmentation: generating segmentation maps to learn structural cues and region partitions.
(4) Image understanding: performing standard vision-language understanding tasks.
The understanding encoder is updated for UMMs with a shared encoder for generation and understanding; otherwise it is frozen.
03

A closer look at the results


      Comparison of UniMRG with other post-training methods for UMMs.
       The green row shows the Gain (↑) over the base model. SFT denotes supervised fine-tuning using only th
Comparison of UniMRG with other post-training methods for UMMs. The green row shows the Gain (↑) over the base model. SFT denotes supervised fine-tuning using only the visual understanding loss.

      Comparison with state-of-the-arts on visual understanding benchmarks. 
       Our model is OpenUni post-trained with UniMRG.
Comparison with state-of-the-arts on visual understanding benchmarks. Our model is OpenUni post-trained with UniMRG.
04

Cite this work

@article{su2026generation,
  title={Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation},
  author={Su, Zihan and Wei, Hongyang and Cen, Kangrui and Wang, Yong and Chen, Guanhua and Yuan, Chun and Chu, Xiangxiang},
  journal={arXiv preprint arXiv:2601.21406},
  year={2026}
}
Enlarged paper figure