arXiv:2601.03198cs.LG2026-01ACL被引 6

构建视觉约束评测集,提升多模态模型对图文指令的准确遵循能力。

Empowering Reliable Visual-Centric Instruction Following in MLLMs

  • 设计含视觉依赖约束的指令,实现图文协同评估
  • 在多个主流模型上提升视觉指令遵循准确率
  • 适合研究多模态推理与指令对齐的学者参考

评估多模态大语言模型(MLLMs)的指令遵循(IF)能力对于严格检验模型输出是否忠实于用户意图至关重要。然而,现有评测基准主要关注文本模态中的口头指令,忽视了语义丰富的视觉模态中隐含的约束,限制了对指令遵循能力的全面分析。为此,我们提出VC-IFEval,一个全新的评测基准及配套系统化构建的数据集,用于在多模态环境下评估MLLMs的指令遵循能力。该基准将视觉依赖性约束系统性地融入指令设计,实现对模型在视觉输入与文本指令双重约束下输出一致性的更严格、细粒度评估。通过在该数据集上微调MLLMs,我们显著提升了视觉指令遵循的准确率与一致性。通过对代表性MLLMs的广泛评估,揭示了当前模型在多模态指令遵循方面的优势与局限。

原文摘要 · Abstract (English)

Evaluating the instruction-following (IF) capabilities of Multimodal Large Language Models (MLLMs) is essential for rigorously assessing how faithfully model outputs adhere to user-specified intentions. Nevertheless, existing benchmarks for evaluating MLLMs' instruction-following capability primarily focus on verbal instructions in the textual modality. These limitations hinder a thorough analysis of instruction-following capabilities, as they overlook the implicit constraints embedded in the semantically rich visual modality. To address this gap, we introduce VC-IFEval, a new benchmark accompanied by a systematically constructed dataset that evaluates MLLMs' instruction-following ability under multimodal settings. Our benchmark systematically incorporates vision-dependent constraints into instruction design, enabling a more rigorous and fine-grained assessment of how well MLLMs align their outputs with both visual input and textual instructions. Furthermore, by fine-tuning MLLMs on our dataset, we achieve substantial gains in visual instruction-following accuracy and adherence. Through extensive evaluation across representative MLLMs, we provide new insights into the strengths and limitations of current models.

多模态指令遵循视觉理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。