arXiv:2509.23879cs.CVcs.AI2025-09EMNLP被引 3

提出新指标评估多模态模型对视觉背景的鲁棒性,发现多数模型易受干扰。

PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications

  • 设计可解释的PCRI指标,量化模型在局部图像块与全图输入下的表现差异
  • 测试19个顶尖模型,仅InternVL2-26B和Qwen2VL-72B表现稳定
  • 为研究人员提供诊断工具,助力开发更可靠的工业级多模态模型

多模态大语言模型在真实场景中的可靠性常因对无关或分散的视觉背景敏感而受损,现有评估指标未能捕捉这一问题。本文提出首个系统且可解释的评分方法——片段上下文鲁棒性指数(Patch Context Robustness Index, PCRI),用于量化模型在不同视觉上下文粒度下的鲁棒性,通过对比局部图像块与全图输入的表现变化实现评估。将PCRI应用于19个领先多模态模型,在15个视觉-语言基准上进行测试,发现多数主流模型仍对背景噪声脆弱,仅有InternVL2-26B和Qwen2VL-72B等少数模型在各项任务中表现出一致的鲁棒性。PCRI分析还揭示了不同模型架构如何处理与融合视觉上下文,为研究者和实践者提供可操作的诊断洞察。该指标支持对上下文鲁棒性的严谨比较,有助于模型选型,并指导未来鲁棒架构与训练策略的开发。

原文摘要 · Abstract (English)

The reliability of Multimodal Large Language Models (MLLMs) in real-world settings is often undermined by sensitivity to irrelevant or distracting visual context, an aspect not captured by existing evaluation metrics. We introduce the \textbf{Patch Context Robustness Index (PCRI)}, the first systematic and interpretable score for quantifying MLLM robustness to variations in visual context granularity, measuring performance changes between localized image patches and full-image input. Applying PCRI to 19 state-of-the-art MLLMs across 15 vision-language benchmarks, we find that most leading models remain brittle to background noise, with only a few, such as InternVL2-26B and Qwen2VL-72B, demonstrating consistent robustness across tasks. PCRI analysis also highlights how different model architectures handle and integrate visual context, offering actionable diagnostic insight for both researchers and practitioners. PCRI enables rigorous comparison of context robustness, supporting principled model selection and guiding the development of future architectures and training strategies for robust, real-world deployment.

多模态鲁棒性评估指标企业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。