arXiv:2512.02014cs.CV2025-12被引 34

TUNA用统一视觉表征打通图文理解与生成,性能更优。

TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models

  • 用级联VAE与表征编码器构建统一视觉空间,支持端到端处理图像视频。
  • 在图文理解、生成和编辑任务上均达最新水平,视频理解提升4.2%。
  • 适合追求统一框架高效多任务的开发者和研究者。

统一多模态模型(UMMs)旨在单一框架内联合完成多模态理解与生成。本文提出TUNA,一种原生统一多模态模型,通过级联变分自编码器(VAE)编码器与表征编码器,构建统一连续视觉表示空间,实现图像与视频的端到端理解与生成。相比以往采用分离表示的模型,TUNA避免了不同编码器间的表征格式不匹配问题,在理解与生成任务上均表现更优。我们还发现,更强的预训练表征编码器在所有多模态任务中均带来性能提升,凸显其重要性。在统一框架下,联合训练理解与生成数据可相互促进而非互相干扰。大量实验表明,TUNA在多模态理解与生成基准上取得当前最优结果,涵盖图像与视频的理解、生成及图像编辑任务,验证了统一表征设计的有效性与可扩展性。

原文摘要 · Abstract (English)

Unified multimodal models (UMMs) aim to jointly perform multimodal understanding and generation within a single framework. We present TUNA, a native UMM that builds a unified continuous visual representation by cascading a VAE encoder with a representation encoder. This unified representation space allows end-to-end processing of images and videos for both understanding and generation tasks. Compared to prior UMMs with decoupled representations, TUNA's unified visual space avoids representation format mismatches introduced by separate encoders, outperforming decoupled alternatives in both understanding and generation. Moreover, we observe that stronger pretrained representation encoders consistently yield better performance across all multimodal tasks, highlighting the importance of the representation encoder. Finally, in this unified setting, jointly training on both understanding and generation data allows the two tasks to benefit from each other rather than interfere. Our extensive experiments on multimodal understanding and generation benchmarks show that TUNA achieves state-of-the-art results in image and video understanding, image and video generation, and image editing, demonstrating the effectiveness and scalability of its unified representation design.

多模态统一表征生成视觉编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。