arXiv:2411.08666cs.CVcs.AI2024-11综述被引 2

系统梳理视觉自回归模型的发展脉络与应用前景

A Survey on Vision Autoregressive Model

  • 将视觉数据转为视觉令牌,统一建模生成与理解任务
  • 覆盖图像、视频、医疗影像等10余类视觉任务
  • 适合关注多模态生成与模型通用性的研究者

自回归模型在自然语言处理中展现出优异的可扩展性、适应性和泛化能力。受此启发,视觉自回归模型近年来被广泛研究,通过将视觉数据表示为视觉令牌,实现对各类视觉任务的自回归建模,涵盖图像生成、视频生成、图像编辑、运动生成、医学影像分析、3D生成、机器人操作及多模态统一生成等。本文系统综述了现有方法,构建了方法分类体系,总结其主要贡献、优势与局限。同时对最新进展进行深入分析,包含跨多个评估数据集的全面基准测试。最后,指出关键挑战并提出未来研究方向,为视觉自回归模型的发展提供路线图。

原文摘要 · Abstract (English)

Autoregressive models have demonstrated great performance in natural language processing (NLP) with impressive scalability, adaptability and generalizability. Inspired by their notable success in NLP field, autoregressive models have been intensively investigated recently for computer vision, which perform next-token predictions by representing visual data as visual tokens and enables autoregressive modelling for a wide range of vision tasks, ranging from visual generation and visual understanding to the very recent multimodal generation that unifies visual generation and understanding with a single autoregressive model. This paper provides a systematic review of vision autoregressive models, including the development of a taxonomy of existing methods and highlighting their major contributions, strengths, and limitations, covering various vision tasks such as image generation, video generation, image editing, motion generation, medical image analysis, 3D generation, robotic manipulation, unified multimodal generation, etc. Besides, we investigate and analyze the latest advancements in autoregressive models, including thorough benchmarking and discussion of existing methods across various evaluation datasets. Finally, we outline key challenges and promising directions for future research, offering a roadmap to guide further advancements in vision autoregressive models.

视觉生成自回归模型多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。