arXiv:2506.09954cs.CVcs.AI2025-06IJCV综述被引 2

综述视觉通用模型的架构、技术与应用前景。

Vision Generalist Model: A Survey

  • 系统梳理视觉通用模型的框架设计与训练技术。
  • 对比多种数据集、任务与评测基准,分析性能差异。
  • 适合关注多模态智能与跨任务泛化的研究者阅读。

近年来,自然语言处理领域取得了通用模型的巨大成功。通用模型是一种在海量数据上训练的通用框架,能够同时处理多种下游任务。受其优异表现的鼓舞,越来越多的研究者开始探索将此类模型应用于计算机视觉任务。然而,视觉任务的输入输出形式更为多样,难以统一表示。本文全面综述了视觉通用模型,深入探讨其特性与能力。首先回顾背景,包括数据集、任务与评测基准;接着分析现有研究提出的框架设计及性能提升技术;为帮助研究者更好地理解该领域,简要介绍相关方向,揭示其关联与潜在协同效应;最后,列举实际应用场景,深入剖析现存挑战,并展望未来研究方向。

原文摘要 · Abstract (English)

Recently, we have witnessed the great success of the generalist model in natural language processing. The generalist model is a general framework trained with massive data and is able to process various downstream tasks simultaneously. Encouraged by their impressive performance, an increasing number of researchers are venturing into the realm of applying these models to computer vision tasks. However, the inputs and outputs of vision tasks are more diverse, and it is difficult to summarize them as a unified representation. In this paper, we provide a comprehensive overview of the vision generalist models, delving into their characteristics and capabilities within the field. First, we review the background, including the datasets, tasks, and benchmarks. Then, we dig into the design of frameworks that have been proposed in existing research, while also introducing the techniques employed to enhance their performance. To better help the researchers comprehend the area, we take a brief excursion into related domains, shedding light on their interconnections and potential synergies. To conclude, we provide some real-world application scenarios, undertake a thorough examination of the persistent challenges, and offer insights into possible directions for future research endeavors.

视觉通用综述多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。