arXiv:2501.02765cs.CVcs.AI2025-01被引 46

综述视觉大语言模型在通用与专用场景的应用现状与挑战

Visual Large Language Models for Generalized and Specialized Applications

  • 系统梳理VLLM在图像、视频、动作等多模态任务中的应用方法
  • 揭示当前VLLM在跨模态理解中的性能瓶颈与伦理风险
  • 适合关注多模态AI前沿与应用落地的研究者参考

视觉语言模型(VLM)已成为学习视觉与语言统一嵌入空间的强大工具。受大型语言模型在推理与多任务能力方面表现的启发,视觉大语言模型(VLLM)正被广泛关注,以构建通用型VLM。尽管在VLLM领域已取得显著进展,相关研究仍相对有限,尤其缺乏从全面应用视角出发的综述,涵盖视觉(图像、视频、深度)、动作和语言模态的通用与专用应用场景。本文聚焦VLLM的多样化应用,分析其使用场景,识别伦理考量与挑战,并探讨未来发展方向。通过整合这些内容,旨在为后续创新与更广泛的应用提供综合指南。论文列表仓库可访问:https://github.com/JackYFL/awesome-VLLMs。

原文摘要 · Abstract (English)

Visual-language models (VLM) have emerged as a powerful tool for learning a unified embedding space for vision and language. Inspired by large language models, which have demonstrated strong reasoning and multi-task capabilities, visual large language models (VLLMs) are gaining increasing attention for building general-purpose VLMs. Despite the significant progress made in VLLMs, the related literature remains limited, particularly from a comprehensive application perspective, encompassing generalized and specialized applications across vision (image, video, depth), action, and language modalities. In this survey, we focus on the diverse applications of VLLMs, examining their using scenarios, identifying ethics consideration and challenges, and discussing future directions for their development. By synthesizing these contents, we aim to provide a comprehensive guide that will pave the way for future innovations and broader applications of VLLMs. The paper list repository is available: https://github.com/JackYFL/awesome-VLLMs.

视觉语言模型多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。