构建多语言抗攻击指令遵循评测集,提升视觉语言模型真实场景可靠性。
MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language Models

- 覆盖中英文任务与多种指令劫持场景,模拟复杂指令约束。
- 含24个子任务、52个子类别,平均每题3.0个约束条件。
- 基于强化学习训练集,显著提升跨语言跨任务泛化能力。
随着视觉语言模型在图像理解、跨模态推理和复杂指令执行方面快速进步,指令遵循能力已成为评估其可靠性与实用性的关键指标。然而,现有多模态指令遵循评测集仍存在语言覆盖有限、对抗安全场景不足的问题,难以有效评估真实世界中的多语言与高安全性场景。为弥补这些缺陷,我们提出MM-IFEval-Pro,一个涵盖中文与英文任务及多样指令劫持案例的多模态指令遵循评测基准。该基准包含4大任务类别、24个子类别,以及8类指令、52个子类别,每条样本平均包含3.0个约束条件,以真实模拟复杂指令情境。我们还构建了一个融合中文与对抗性指令的强化学习训练集,显著提升了模型在MM-IFEval-Pro上的表现,并在主流多模态基准上实现良好迁移,展现出强跨任务与跨语言泛化能力。
原文摘要 · Abstract (English)
As vision-language models (VLMs) rapidly advance in image understanding, cross-modal reasoning, and complex instruction execution, instruction-following capability has become a key indicator of their reliability and practicality. However, existing multimodal instruction-following benchmarks still suffer from limited language coverage and insufficient adversarial safety scenarios, making them inadequate for evaluating real-world multilingual and safety-sensitive settings. To address these gaps, we present MM-IFEval-Pro, a multimodal instruction-following benchmark covering Chinese and English tasks as well as diverse instruction hijacking cases. MM-IFEval-Pro includes 4 major task categories and 24 subcategories and 8 instruction categories with 52 subcategories, with each sample containing an average of 3.0 constraints to realistically simulate complex instruction scenarios. We further construct a reinforcement-learning training set enriched with Chinese and adversarial instructions, which significantly improves model performance on MM-IFEval-Pro and transfers effectively to other mainstream multimodal benchmarks, demonstrating strong cross-task and cross-language generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。