arXiv:2410.22217cs.CV2024-10综述被引 13

探索视觉大模型中统一理解与生成的新范式

Towards Unifying Understanding and Generation in the Era of Vision Foundation Models: A Survey from the Autoregression Perspective

  • 从自回归视角系统梳理视觉大模型发展
  • 提出统一理解和生成的视觉建模新趋势
  • 适合关注多任务视觉模型的研究者参考

大型语言模型中的自回归机制通过将所有语言任务统一为下一个词预测,展现出惊人可扩展性。近期,学界开始尝试将这一成功经验拓展至视觉基础模型。本文综述了自回归视觉基础模型的最新进展,并探讨未来方向。首先,指出下一代视觉基础模型的发展趋势:统一视觉任务的理解与生成能力。接着分析现有模型的局限性,正式定义自回归机制及其优势。随后,根据视觉分词器与自回归骨干网络对模型进行分类。最后讨论若干有前景的研究挑战与方向。据我们所知,这是首篇全面总结自回归视觉基础模型在统一理解与生成趋势下的综述。相关资源可访问 https://github.com/EmmaSRH/ARVFM。

原文摘要 · Abstract (English)

Autoregression in large language models (LLMs) has shown impressive scalability by unifying all language tasks into the next token prediction paradigm. Recently, there is a growing interest in extending this success to vision foundation models. In this survey, we review the recent advances and discuss future directions for autoregressive vision foundation models. First, we present the trend for next generation of vision foundation models, i.e., unifying both understanding and generation in vision tasks. We then analyze the limitations of existing vision foundation models, and present a formal definition of autoregression with its advantages. Later, we categorize autoregressive vision foundation models from their vision tokenizers and autoregression backbones. Finally, we discuss several promising research challenges and directions. To the best of our knowledge, this is the first survey to comprehensively summarize autoregressive vision foundation models under the trend of unifying understanding and generation. A collection of related resources is available at https://github.com/EmmaSRH/ARVFM.

视觉模型自回归统一任务大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。