arXiv:2411.05902cs.CVcs.CL2024-11中稿 · TMLR综述被引 60

系统梳理视觉领域自回归模型的演进与应用,涵盖图像、视频、3D生成等方向。

Autoregressive Models in Vision: A Survey

  • 按像素、标记、尺度三类表示策略分类视觉自回归模型
  • 梳理250篇相关论文,覆盖图像到3D医学生成的多模态应用
  • 适合想了解生成模型在视觉中发展脉络的研究者参考

自回归模型在自然语言处理中已取得巨大成功。近年来,该类模型在计算机视觉领域也成为重要研究方向,能够生成高质量视觉内容。虽然自然语言处理中的自回归模型通常作用于子词标记,但计算机视觉的表示策略可分像素级、标记级或尺度级,反映了视觉数据相较于语言序列的多样性和层次性。本文全面综述了自回归模型在视觉领域的研究进展。为便于不同背景的研究者理解,先介绍视觉中的序列表示与建模基础。随后根据表示策略将视觉自回归模型分为三类:基于像素、基于标记和基于尺度的模型。进一步探讨其与其他生成模型的关联。同时,从图像生成、视频生成、3D生成及多模态生成等多个维度进行细分。还介绍了其在具身智能、3D医学人工智能等新兴领域的应用。全文包含约250篇相关文献,并指出当前挑战与潜在研究方向。相关论文资料已整理至GitHub仓库:https://github.com/ChaofanTao/Autoregressive-Models-in-Vision-Survey。

原文摘要 · Abstract (English)

Autoregressive modeling has been a huge success in the field of natural language processing (NLP). Recently, autoregressive models have emerged as a significant area of focus in computer vision, where they excel in producing high-quality visual content. Autoregressive models in NLP typically operate on subword tokens. However, the representation strategy in computer vision can vary in different levels, i.e., pixel-level, token-level, or scale-level, reflecting the diverse and hierarchical nature of visual data compared to the sequential structure of language. This survey comprehensively examines the literature on autoregressive models applied to vision. To improve readability for researchers from diverse research backgrounds, we start with preliminary sequence representation and modeling in vision. Next, we divide the fundamental frameworks of visual autoregressive models into three general sub-categories, including pixel-based, token-based, and scale-based models based on the representation strategy. We then explore the interconnections between autoregressive models and other generative models. Furthermore, we present a multifaceted categorization of autoregressive models in computer vision, including image generation, video generation, 3D generation, and multimodal generation. We also elaborate on their applications in diverse domains, including emerging domains such as embodied AI and 3D medical AI, with about 250 related references. Finally, we highlight the current challenges to autoregressive models in vision with suggestions about potential research directions. We have also set up a Github repository to organize the papers included in this survey at: https://github.com/ChaofanTao/Autoregressive-Models-in-Vision-Survey.

自回归模型视觉生成综述多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。