评测多模态大模型在视频理解中遵循复杂指令的能力。
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

- 构建四类指令模板,覆盖32种任务和39类约束。
- 1.5K样本通过半自动流程生成,兼顾效率与准确。
- 现有模型对多约束、语义类或条件分支指令仍表现不佳。
多模态大语言模型(MLLMs)在视频理解任务中表现优异,但其在该领域遵循用户指令的能力尚未充分探索。真实场景中的视频理解不仅要求正确解析内容,还需满足多样化的用户指定约束。现有基准主要关注任务准确性,忽视指令遵循能力。为此,我们提出Video-IFBench,一个全面评估视频理解中指令遵循能力的基准,要求模型满足基于视觉与音频内容的多样化约束。我们设计了包含单任务、多任务、选择与嵌套指令的四类模板,涵盖32种任务类型与39个手动设计的约束类别,涉及语义与格式要求。为降低标注成本,我们构建了结合MLLM、程序化处理与人工验证的半自动数据生成流程,最终获得1.5K高质量样本。我们对超过20个近期的MLLM进行了大规模评估,结果表明当前模型在应对多约束、语义类或具有复杂条件结构的指令时仍面临显著挑战,尤其在根据视频内容选择正确执行路径方面表现不足。我们希望本工作能推动视频理解中指令遵循能力的研究进展。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understanding, where models must satisfy diverse user-specified constraints, including those grounded in visual and audio content. We develop an instruction taxonomy with four templates, including single-task, multi-task, selection, and nested instructions, covering 32 task types and 39 manually designed constraint categories spanning both semantic and format requirements. To reduce annotation cost, we build a semi-automatic data construction pipeline that combines MLLMs, programmatic processing, and human verification, resulting in 1.5K samples. We conduct a large-scale evaluation of more than 20 recent MLLMs and show that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content. We hope our work will facilitate future research on instruction following in video understanding scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。