arXiv:2512.13752cs.CVcs.AI2025-12被引 6

STAR通过分阶段训练提升多模态模型的生成与理解能力。

STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning

  • 分理解、生成、编辑三阶段,逐步叠加自回归模块。
  • 在GenEval、DPG-Bench等任务上达最新水平,最高得分87.44。
  • 适合需要统一生成与理解能力的多模态研究者使用。

多模态大语言模型(MLLMs)在推动通用人工智能发展方面发挥关键作用。然而,由于优化冲突和性能权衡,实现统一的多模态理解和生成仍具挑战。为有效提升生成性能并保留现有理解能力,我们提出STAR:一种用于任务渐进式统一多模态学习的堆叠自回归方案。该方法将多模态学习分解为理解、生成和编辑三个阶段,通过冻结基础自回归(AR)模型参数,并逐步堆叠同构的AR模块,避免跨任务干扰的同时扩展模型能力。同时,引入高容量向量量化(VQ)以增强图像表征粒度,并采用隐式推理机制,在复杂条件下提升生成质量。实验表明,STAR在GenEval(0.91)、DPG-Bench(87.44)和ImgEdit(4.34)上均达到当前最优表现,验证了其在统一多模态学习中的有效性。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) play a pivotal role in advancing the quest for general artificial intelligence. However, achieving unified target for multimodal understanding and generation remains challenging due to optimization conflicts and performance trade-offs. To effectively enhance generative performance while preserving existing comprehension capabilities, we introduce STAR: a STacked AutoRegressive scheme for task-progressive unified multimodal learning. This approach decomposes multimodal learning into multiple stages: understanding, generation, and editing. By freezing the parameters of the fundamental autoregressive (AR) model and progressively stacking isomorphic AR modules, it avoids cross-task interference while expanding the model's capabilities. Concurrently, we introduce a high-capacity VQ to enhance the granularity of image representations and employ an implicit reasoning mechanism to improve generation quality under complex conditions. Experiments demonstrate that STAR achieves state-of-the-art performance on GenEval (0.91), DPG-Bench (87.44), and ImgEdit (4.34), validating its efficacy for unified multimodal learning.

多模态自回归生成统一学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。