让视觉模型一边生成一边理解,提升识别精度和生成质量。
Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation
- 用重建、深度、分割三种方式生成图像表示,辅助理解
- 在多个模型上验证,显著减少幻觉并提升空间感知能力
- 适合想同时优化理解和生成的多模态研究者
统一多模态模型(UMMs)将视觉理解与生成整合于同一框架中。尽管现有后训练方法已利用理解来增强生成,但如何通过生成反向提升理解仍鲜有探索。本文提出架构无关的UniMRG方法,通过引入辅助生成任务来增强理解。具体地,在标准视觉理解目标外,训练模型生成输入图像的像素级重建、深度图(几何)和分割图(结构)等多类内在表示。这些多样化表示捕捉了外观、空间关系与结构布局的互补信息,使模型对视觉输入形成更深层、全面的理解。在多种不同架构的UMM上进行的大量实验表明,该方法显著提升了细粒度感知能力,减少了幻觉现象,并改善了空间理解,同时增强了生成性能。
原文摘要 · Abstract (English)
Unified Multimodal Models (UMMs) integrate both visual understanding and generation within a single framework. Their ultimate aspiration is to create a cycle where understanding and generation mutually reinforce each other. While recent post-training methods have successfully leveraged understanding to enhance generation, the reverse direction of utilizing generation to improve understanding remains largely unexplored. In this work, we propose UniMRG (Unified Multi-Representation Generation), a simple yet effective architecture-agnostic post-training method. UniMRG enhances the understanding capabilities of UMMs by incorporating auxiliary generation tasks. Specifically, we train UMMs to generate multiple intrinsic representations of input images, namely pixel (reconstruction), depth (geometry), and segmentation (structure), alongside standard visual understanding objectives. By synthesizing these diverse representations, UMMs capture complementary information regarding appearance, spatial relations, and structural layout. Consequently, UMMs develop a deeper and more comprehensive understanding of visual inputs. Extensive experiments across diverse UMM architectures demonstrate that our method notably enhances fine-grained perception, reduces hallucinations, and improves spatial understanding, while simultaneously boosting generation capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。