分析2.6万篇顶会论文,揭示视觉语言模型的三大趋势
Vision Language Models: A Survey of 26K Papers
- 基于标题摘要自动打标签,量化研究热点演变
- 多模态大模型崛起,生成技术聚焦可控性与效率
- 适合关注AI研究动态的学者和从业者
本文对2023至2025年CVPR、ICLR、NeurIPS收录的26,104篇论文进行透明可复现的分析。通过标准化标题与摘要,结合人工构建词典,为每篇论文分配最多35个主题标签,挖掘任务、架构、训练方式、目标函数、数据集及跨模态共现等细粒度线索。分析显示三大宏观趋势:(1)多模态视觉-语言-大模型显著增长,经典感知任务逐渐被指令跟随与多步推理重构;(2)生成方法持续扩展,扩散模型研究集中于可控性、蒸馏与速度提升;(3)3D与视频研究保持活跃,从NeRF转向高斯点云表示,日益强调人类与智能体为中心的理解。在视觉语言模型领域,参数高效适配(如提示/适配器/LoRA)和轻量级视觉-语言桥梁占主导;训练范式由自建编码器转向对强基座模型的指令微调与精调;对比学习目标退居其次,交叉熵/排序损失与蒸馏更受青睐。跨会议比较显示,CVPR 3D相关论文更多,ICLR 视觉语言模型占比最高,效率与鲁棒性等可靠性议题在各领域广泛传播。我们公开词典与方法论以支持审计与拓展。局限包括词典召回率与仅基于摘要的分析范围,但纵向信号在不同会议与年份间具有一致性。
原文摘要 · Abstract (English)
We present a transparent, reproducible measurement of research trends across 26,104 accepted papers from CVPR, ICLR, and NeurIPS spanning 2023-2025. Titles and abstracts are normalized, phrase-protected, and matched against a hand-crafted lexicon to assign up to 35 topical labels and mine fine-grained cues about tasks, architectures, training regimes, objectives, datasets, and co-mentioned modalities. The analysis quantifies three macro shifts: (1) a sharp rise of multimodal vision-language-LLM work, which increasingly reframes classic perception as instruction following and multi-step reasoning; (2) steady expansion of generative methods, with diffusion research consolidating around controllability, distillation, and speed; and (3) resilient 3D and video activity, with composition moving from NeRFs to Gaussian splatting and a growing emphasis on human- and agent-centric understanding. Within VLMs, parameter-efficient adaptation like prompting/adapters/LoRA and lightweight vision-language bridges dominate; training practice shifts from building encoders from scratch to instruction tuning and finetuning strong backbones; contrastive objectives recede relative to cross-entropy/ranking and distillation. Cross-venue comparisons show CVPR has a stronger 3D footprint and ICLR the highest VLM share, while reliability themes such as efficiency or robustness diffuse across areas. We release the lexicon and methodology to enable auditing and extension. Limitations include lexicon recall and abstract-only scope, but the longitudinal signals are consistent across venues and years.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。