arXiv:2412.08158cs.CVcs.CL2024-12综述被引 1

大模型如何提升视觉语言任务表现的系统梳理

How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey

  • 从经典挑战出发,分析预训练模型如何重构视觉语言任务范式
  • 总结多类任务中预训练模型带来的性能提升与新方法演进
  • 揭示模型局限性并指明未来研究方向,适合关注AI跨模态进展者

视觉语言任务(如图像描述生成、视觉问答和视觉常识推理)是人工智能的重要研究方向,持续吸引学术界关注。尽管整体性能不断提升,经典挑战依然存在,制约该领域发展。近年来,预训练模型的兴起推动了视觉语言任务的研究进步。凭借海量训练数据与参数规模,预训练模型在多个下游任务中展现出卓越性能。受其强大能力启发,新的解决范式涌现,已成为当前主流研究方向,并引发快速进展。本文系统综述预训练模型如何赋能视觉语言任务:首先回顾该领域的主要挑战及预训练时代前解决方案的局限;其次总结近期利用预训练模型应对挑战的最新进展;最后分析预训练模型固有局限带来的潜在风险,探讨可能解决方案,尝试为未来研究提供指引。

原文摘要 · Abstract (English)

The exploration of various vision-language tasks, such as visual captioning, visual question answering, and visual commonsense reasoning, is an important area in artificial intelligence and continuously attracts the research community's attention. Despite the improvements in overall performance, classic challenges still exist in vision-language tasks and hinder the development of this area. In recent years, the rise of pre-trained models is driving the research on vision-language tasks. Thanks to the massive scale of training data and model parameters, pre-trained models have exhibited excellent performance in numerous downstream tasks. Inspired by the powerful capabilities of pre-trained models, new paradigms have emerged to solve the classic challenges. Such methods have become mainstream in current research with increasing attention and rapid advances. In this paper, we present a comprehensive overview of how vision-language tasks benefit from pre-trained models. First, we review several main challenges in vision-language tasks and discuss the limitations of previous solutions before the era of pre-training. Next, we summarize the recent advances in incorporating pre-trained models to address the challenges in vision-language tasks. Finally, we analyze the potential risks associated with the inherent limitations of pre-trained models and discuss possible solutions, attempting to provide future research directions.

视觉语言预训练模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。