arXiv:2510.18726cs.CV2025-10被引 6

评测视频描述模型能否听懂用户指令,发现开源模型已逼近闭源水平。

IF-VidCap: Can Video Caption Models Follow Instructions?

  • 构建新基准IF-VidCap,从格式和内容双维度评估指令遵循能力。
  • 20多个模型测试显示,顶尖开源模型在指令跟随上接近闭源模型。
  • 专门做密集描述的模型反而不擅长复杂指令,需兼顾描述与指令理解。

尽管多模态大语言模型在视频字幕生成上表现出色,但实际应用需要符合用户特定指令的描述,而非冗长无约束的内容。现有基准主要评估描述的全面性,却忽视了指令遵循能力。为此,我们提出IF-VidCap,一个全新的可控视频字幕评估基准,包含1,400个高质量样本。不同于现有视频字幕或通用指令遵循基准,IF-VidCap采用系统化框架,从格式正确性和内容正确性两个维度评估。对20多个主流模型的全面评估显示:尽管闭源模型仍占优势,但差距正在缩小,顶级开源模型已实现近似对齐。此外,我们发现专注于密集字幕的模型在复杂指令下表现不如通用型MLLMs,提示未来研究需同步提升描述丰富性与指令遵循精度。

原文摘要 · Abstract (English)

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained descriptions. Current benchmarks, however, primarily assess descriptive comprehensiveness while largely overlooking instruction-following capabilities. To address this gap, we introduce IF-VidCap, a new benchmark for evaluating controllable video captioning, which contains 1,400 high-quality samples. Distinct from existing video captioning or general instruction-following benchmarks, IF-VidCap incorporates a systematic framework that assesses captions on two dimensions: format correctness and content correctness. Our comprehensive evaluation of over 20 prominent models reveals a nuanced landscape: despite the continued dominance of proprietary models, the performance gap is closing, with top-tier open-source solutions now achieving near-parity. Furthermore, we find that models specialized for dense captioning underperform general-purpose MLLMs on complex instructions, indicating that future work should simultaneously advance both descriptive richness and instruction-following fidelity.

视频生成指令遵循多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。