arXiv:2606.04588cs.CL2026-06

评测视频理解模型对复杂指令的遵循能力,发现多数模型难以同时满足多重约束。

VCIFBench: Evaluating Complex Instruction Following for Video Understanding

论文配图:VCIFBench: Evaluating Complex Instruction Following for Video Understanding
图 1 · 摘自论文原文
  • 构建包含内容、格式等多维度约束的指令,模拟真实复杂场景。
  • 10个模型均难以同时满足全部约束,最优仅达73%正确率。
  • 揭示模型常忽略全局冲突,只执行部分可满足指令。

多模态大语言模型在视频理解领域进展迅速,但现有评测基准多依赖简单提示,缺乏对模型是否能遵守明确输出约束的验证。本文提出VCIFBench,一个用于评估视频理解中复杂指令遵循能力的基准。该基准从适配基准和直接视频引导的提示中构建高约束性指令,涵盖内容、格式、风格与结构要求,并采用混合验证流程评估模型输出。基准包含306个可满足测试指令、540个DPO训练样本及100项诊断集,用于检验模型识别指令冲突的能力。对10个MLLMs的实验表明,联合满足多种约束仍具挑战性;偏好优化提升了两类模型的指令遵循能力,而Conflict-100结果显示,模型通常执行看似可满足的子集,而非识别全局不相容性。

原文摘要 · Abstract (English)

Multimodal large language models have made rapid progress in video understanding, yet existing benchmarks largely rely on simple prompts and provide limited evidence about whether models can satisfy explicit output constraints. We introduce VCIFBench, a benchmark for evaluating complex instruction following in video understanding. VCIFBench constructs constraint-rich instructions from both benchmark-adapted and directly video-grounded prompts, covering content, format, style, and structure requirements, and evaluates model outputs with a hybrid verification pipeline. The benchmark contains 306 satisfiable test instructions, 540 DPO training instances, and a 100-item diagnostic set for evaluating whether models can recognize instruction conflicts. Experiments on 10 MLLMs show that joint constraint satisfaction remains challenging. Preference optimization improves instruction following for two model families, while Conflict-100 reveals that models usually execute a satisfiable-looking subset instead of detecting global incompatibility.

视频理解指令遵循评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。