arXiv:2603.27662cs.CV2026-03

评测8个开源视频大模型在新闻视频自动生成字幕的表现

A Benchmarking Methodology to Assess Open-Source Video Large Language Models in Automatic Captioning of News Videos

  • 用双数据集对比8个开源模型,涵盖智利和英国新闻片段
  • 新指标TFS/EFS更准确评估主题与实体保留,优于传统指标
  • Gemma~3表现最佳,适合需高精度字幕的新闻自动化场景

新闻视频是电视台和在线平台最常见内容类型之一,但生成文本描述以辅助索引与检索仍主要依赖人工。视频大语言模型(VidLLMs)有望实现该任务的自动化,但在新闻领域尚缺乏全面评估。本文对8个最先进的开源VidLLMs在自动新闻视频字幕生成任务中进行比较研究,使用两个互补的基准数据集:约1,345段的智利电视新闻语料库和9,838段的BBC新闻语料库。评估采用词汇度量(METEOR、ROUGE-L)、语义度量(BERTScore、CLIPScore、Text Similarity、Mean Reciprocal Rank),以及本文提出的两个新保真度指标:主题保真度评分(TFS)和实体保真度评分(EFS)。分析表明,标准度量因依赖表面形式、对静态帧不敏感及功能词膨胀,在新闻视频字幕评估中区分能力有限。TFS与EFS通过直接评估生成字幕的主题结构保持和命名实体覆盖情况,弥补了这一不足。结果表明,Gemma~3在两个数据集及多数评估维度上表现最佳,Qwen-VL为稳定第二。

原文摘要 · Abstract (English)

News videos are among the most prevalent content types produced by television stations and online streaming platforms, yet generating textual descriptions to facilitate indexing and retrieval largely remains a manual process. Video Large Language Models (VidLLMs) offer significant potential to automate this task, but a comprehensive evaluation in the news domain is still lacking. This work presents a comparative study of eight state-of-the-art open-source VidLLMs for automatic news video captioning, evaluated on two complementary benchmark datasets: a Chilean TV news corpus (approximately 1,345 clips) and a BBC News corpus (9,838 clips). We employ lexical metrics (METEOR, ROUGE-L), semantic metrics (BERTScore, CLIPScore, Text Similarity, Mean Reciprocal Rank), and two novel fidelity metrics proposed in this work: the Thematic Fidelity Score (TFS) and Entity Fidelity Score (EFS). Our analysis reveals that standard metrics exhibit limited discriminative power for news video captioning due to surface-form dependence, static-frame insensitivity, and function-word inflation. TFS and EFS address these gaps by directly assessing thematic structure preservation and named-entity coverage in the generated captions. Results show that Gemma~3 achieves the highest overall performance across both datasets and most evaluation dimensions, with Qwen-VL as a consistent runner-up.

视频生成大模型评测新闻字幕多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。