arXiv:2601.06566cs.CVcs.AI2026-01被引 1

融合多模型提升视频理解,实现精准图文问答与生成。

QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models

  • 三模型协同:关键帧提取+多模态大模型+语言大模型
  • 图文问答准确率提升48.9%,视频描述质量提高44.2%
  • 适合需本地部署的智能视频分析场景

本文提出QCaption,一种融合三种模型的视频字幕生成与问答新框架:关键帧提取、用于图文分析的大规模多模态模型(LMM),以及用于文本分析的大规模语言模型(LLM)。该方法实现文本、图像与视频的联合分析,在视频字幕生成和问答任务上分别取得最高44.2%和48.9%的性能提升。实验验证了该融合策略的有效性,同时通过消融研究评估了LLM在融合中的作用。此外,论文还提出了若干新的视频字幕生成方法,并与QCaption及现有方法进行对比。结果表明,模型融合可显著推动视频分析技术发展。

原文摘要 · Abstract (English)

This paper introduces QCaption, a novel video captioning and Q&A pipeline that enhances video analytics by fusing three models: key frame extraction, a Large Multimodal Model (LMM) for image-text analysis, and a Large Language Model (LLM) for text analysis. This approach enables integrated analysis of text, images, and video, achieving performance improvements over existing video captioning and Q&A models; all while remaining fully self-contained, adept for on-premises deployment. Experimental results using QCaption demonstrated up to 44.2% and 48.9% improvements in video captioning and Q&A tasks, respectively. Ablation studies were also performed to assess the role of LLM on the fusion on the results. Moreover, the paper proposes and evaluates additional video captioning approaches, benchmarking them against QCaption and existing methodologies. QCaption demonstrate the potential of adopting a model fusion approach in advancing video analytics.

视频理解多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。