arXiv:2509.16538cs.CVcs.CL2025-09ACL被引 2

无需参考文本,就能精准评估视频字幕事实准确性的轻量级模型

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis

  • 基于可控错误生成的训练框架,提升评估可解释性
  • 在多个基准上与人工评价相关性达最新水平
  • 适合需要真实世界视频字幕质量评估的研究者

我们提出VC-Inspector,一种轻量级、开源的大规模多模态模型(LMM),用于无参考文本的视频字幕事实性评估。与现有指标在上下文处理能力有限、事实判断较弱或依赖专有服务的问题不同,VC-Inspector提供可复现且注重事实的评估方式,其结果与人类判断高度一致。为实现稳健训练和可解释评估,我们构建了系统化框架,生成带有可控事实错误的字幕,并附带分级质量评分和解释性标注。实验表明,VC-Inspector在多个领域(如VATEX-Eval、Flickr8K-Expert和Flickr8K-CF)均达到当前最优的人类评价相关性,揭示了字幕改进的潜力。项目页面见 https://dipta007.github.io/VC-Inspector

原文摘要 · Abstract (English)

We propose VC-Inspector, a lightweight, open-source large multimodal model (LMM) for reference-free evaluation of video captions, with a focus on factual accuracy. Unlike existing metrics that suffer from limited context handling, weak factuality assessment, or reliance on proprietary services, VC-Inspector offers a reproducible and fact-aware alternative that aligns closely with human judgments. To enable robust training and interpretable evaluation, we introduce a systematic framework for generating captions with controllable factual errors, paired with graded quality scores and explanatory annotations. Experiments demonstrate that VC-Inspector achieves state-of-the-art correlation with human judgments, generalizing across diverse domains (e.g., VATEX-Eval, Flickr8K-Expert, and Flickr8K-CF benchmarks) and revealing the potential for caption improvement. Project page is available at https://dipta007.github.io/VC-Inspector

视频字幕事实评估无参考评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。