Generation Enhances Understanding
in Unified Multimodal Models via
Multi-Representation Generation
Pixel, depth, and segmentation.
Generation as a path to richer visual understanding.
Pixel
Appearance
Image reconstructionDepth
Geometry
Spatial relationsSegmentation
Structure
Region partitionsGeneration → understanding01 / The idea
Pixel. Depth. Segmentation.
Multiple representations. A deeper understanding.

Generation helps models
understand what they see.
Unified multimodal models can both understand and generate images. Yet using generation to improve understanding remains less explored.
UniMRG trains models to generate pixels, depth maps, and segmentation maps alongside visual understanding tasks. These complementary signals improve fine-grained perception, reduce hallucinations, and strengthen spatial understanding.
Read the full abstract ↗One model, complementary representations

(1) Image reconstruction: reconstructing the input image to enhance generation capabilities.
(2) Image-to-depth: generating depth maps to learn geometric cues and spatial relations.
(3) Image-to-segmentation: generating segmentation maps to learn structural cues and region partitions.
(4) Image understanding: performing standard vision-language understanding tasks.
The understanding encoder is updated for UMMs with a shared encoder for generation and understanding; otherwise it is frozen.
A closer look at the results




Cite this work
@article{su2026generation,
title={Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation},
author={Su, Zihan and Wei, Hongyang and Cen, Kangrui and Wang, Yong and Chen, Guanhua and Yuan, Chun and Chu, Xiangxiang},
journal={arXiv preprint arXiv:2601.21406},
year={2026}
}