探索视觉如何驱动通用智能,为构建通用人工智能提供新路径。
Visual General Intelligence: A White Paper
- 以视觉为核心,探讨智能从图像视频中涌现的可能
- 强调多模态协同与视觉输入在智能系统中的基础作用
- 适合关注通用人工智能与计算机视觉融合的研究者
本文从视觉中心视角重新思考智能的本质,探讨视觉经验与学习是否能成为通向通用人工智能(AGI)的路径。语言领域自Transformer架构出现后,GPT系列通过大规模文本自回归建模实现了对未见任务的迁移能力,这引发了一个自然问题:图像、视频和几何等视觉模态能否催生类似的能力?本文汇聚多方观点,讨论视觉智能(VGI)作为通向AGI的潜在路径,不提供单一定义,而是明确计算机视觉在AGI时代应遵循的原则、核心输入模态、评估基准、学习范式,以及视觉与其他模态(如语言)的关系。
原文摘要 · Abstract (English)
This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we discuss whether visual intelligence can serve as a pathway toward AGI, referred to in this paper as visual general intelligence (VGI), by bringing together contributors from diverse standpoints and affiliations. Our aim is not to offer a single definition of visual intelligence, but to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and the relationship between vision, when taken as the core, and other modalities such as language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。