综述2026年前沿多模态大模型的演进、评测与挑战
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
- 梳理从双塔模型到统一感知-生成-行动模型的架构发展
- 评测体系扩展至长视频、具身任务和奖励模型评判
- 聚焦幻觉、安全、对齐等关键难题,适合研究者参考
多模态大模型(VLMs)已从对比图像文本编码器和适配器型助手,演变为支持长上下文推理、代理工作流及日益统一的感知-生成-行动闭环的原生多模态基础模型。自2025年起,前沿集中于少数大型预训练家族,包括GPT、Gemini、Claude、Grok、Qwen、Gemma、DeepSeek、Kimi和MiniMax。评估重点从短文本视觉问答转向空间、时间、具身及校准敏感性基准。本文综述至2026年的VLM进展,重点关注三方面:第一,架构演化路径:从对比/桥接双塔模型,到以大语言模型为骨干的适配器模型、原生多模态输入模型、全模态统一输入输出模型,再到新兴的世界-动作模型;第二,评测体系扩展:涵盖长视频推理、具身评估、奖励模型判断及领域特定决策;第三,后训练角色增强:包括监督微调、偏好优化以及类似GRPO的多模态对齐强化学习方法。本文不穷举所有发布模型,而是聚焦代表性前沿家族、有影响力基准与当前仍存在的核心挑战——幻觉、安全、效率、数据质量与多模态对齐。
原文摘要 · Abstract (English)
Large vision-language models (VLMs) have evolved rapidly from contrastive image-text encoders and adapter-based assistants into natively multimodal foundation models that support long-context reasoning, agentic workflows, and, increasingly, unified perception-generation-action loops. Since 2025, the frontier has consolidated around a small number of large pretraining families, including GPT, Gemini, Claude, Grok, Qwen, Gemma, DeepSeek, Kimi, and MiniMax, while evaluation has shifted from short-form visual question answering toward spatial, temporal, embodied, and calibration-sensitive benchmarks. In this survey, we provide an updated overview of VLMs through 2026 with an emphasis on three developments. First, we trace the architectural evolution from contrastive or bridged two-tower models to LLM-backbone adapter models, native multimodal-input models, omni-modal unified input/output models, and emerging world-action models. Second, we summarize how benchmarks expanded from classical OCR, VQA, and chart understanding toward long-video reasoning, embodied evaluation, reward-model judging, and domain-specific decision making. Third, we review the growing role of post-training, including supervised fine-tuning, preference optimization, and reinforcement learning methods such as GRPO-style multimodal alignment. Rather than exhaustively listing every release, we focus on representative frontier families, influential benchmarks, and the major open challenges that remain in hallucination, safety, efficiency, data quality, and multimodal alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。