测试多模态大模型对系统指令的遵守程度,发现越复杂的视觉指令越难执行。
Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System Messages

- 构建新基准VSysBench,分类22类指令并配对反例测试模型服从性
- 16个模型中系统指令使基础任务准确率显著下降,视觉指令最难遵守
- 开源模型在用户冲突下服从性崩溃,闭源领先模型保持稳定
多模态大模型在实际部署中越来越依赖系统消息来控制行为。然而现有评估基准要么仅针对文本约束,要么将指令嵌入用户输入,导致多模态情境下对系统消息的遵守程度缺乏测量;同时未明确遵守指令是否以牺牲基础视觉语言能力为代价。我们提出VSysBench,基于MMVet-v2构建,将约束分为5大类22小类,涵盖从视觉语境中的文本指令到完全基于视觉的指令,并为每类配备不一致的对照样本以测试指令优先级。通过联合满意度率(JSR)和跨约束敏感度(CCS)两个维度评估模型响应。在16个MLLM上测试发现:施加系统消息会显著降低基础任务准确率;开放权重模型在用户冲突下服从性崩溃,而顶级闭源模型仍保持稳定;所有模型中,视觉基底约束是最难遵守的一类。
原文摘要 · Abstract (English)
Production deployments of Multimodal Large Language Models (MLLMs) increasingly rely on system messages to govern model behavior. Yet existing benchmarks either evaluate constraints in text only or embed them into the user turn, leaving system-message adherence in multimodal contexts largely unmeasured; they also leave open whether compliance comes at the cost of foundational vision-language capabilities. We introduce VSysBench, a benchmark built on MMVet-v2 that organizes constraints into 5 main categories and 22 sub-categories, ranging from textual directives in visual contexts to fully vision-grounded ones, each paired with a misaligned counterpart that stress-tests the instructional hierarchy. VSysBench scores each response jointly along two axes, constraint compliance and answer correctness, via the Joint Satisfaction Rate (JSR) and Cross-Constraint Sensitivity (CCS). Across 16 MLLMs, we find that imposing system messages substantially erodes base task accuracy, that compliance collapses under user conflict for open-weight models while remaining stable for top proprietary ones, and that vision-grounded constraints are the hardest category for every model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。