检测多模态大模型对输入顺序的敏感性,发现所有模型都受顺序影响。
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

- 设计五维审计框架,测试不同模态顺序对答案的影响。
- 18个模型翻转率24%-50%,最佳模型仍有13.4%错误。
- 提示词调整无法通用解决顺序问题,需从训练和架构入手。
多模态大语言模型(MLLMs)的标准评测仅使用单一输入顺序,忽略了顺序无关性这一基础可靠性要求。本文提出Facet-Probe,对18个前沿及开源的MLLMs进行五维审计(选项、证据块、文档排序、图像集、混合模态顺序)。基于贝叶斯项目反应模型分离顺序噪声与各维度偏倚,并通过同顺序控制估计解码器随机性基线。结果显示,18个模型均非顺序不变:各维度面板均值翻转率在24%-50%之间。以温度0的Gemini为同输入控制,验证单元中仍存在显著的顺序超额翻转。模型能力可预测但无法消除翻转,最优模型在13.4%的测试中出错。针对Gemini的缓解实验表明,无需训练的提示调整具有模态特异性且无法跨模态迁移。结果表明,仅靠提示级优化难以实现普遍的顺序鲁棒性,亟需训练与架构层面的新方法。建议将跨顺序翻转率作为MLLM标准报告维度。
原文摘要 · Abstract (English)
Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines. We introduce Facet-Probe, a five-facet audit (option, evidence-chunk, document-rank, image-set, and mixed-modality ordering) of 18 frontier and open-weight MLLMs. A Bayesian item-response model separates ordering noise from per-facet bias, and a same-ordering control estimates the decoder-stochastic floor for observed flips. We find that none of the 18 MLLMs we audit are order-invariant: screened per-facet panel-mean flip rates span 24-50%. A Gemini same-ordering control at temperature 0 estimates a substantial ordering excess over a same-input decoder-noise floor in verified cells. Capability predicts but does not eliminate flips; the best model still flips on 13.4% of trials. In our Gemini mitigation tests, training-free prompt changes are modality-conditional and do not transfer from text to visual reasoning. These results suggest that prompt-level mitigation alone is unlikely to provide general order robustness, motivating future work on training-time and architectural approaches. We propose cross-ordering flip rate as a standard reporting axis for MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。