arXiv:2608.07742cs.CV2026-08

评测科学视觉语言模型在图像退化下的推理稳定性,发现模型易因模糊等干扰失效。

BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning

论文配图:BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning
图 1 · 摘自论文原文
  • 通过渐进式图像退化测试模型鲁棒性,模拟真实场景中的画质下降。
  • 提出两项新指标,量化模型在视觉退化下推理性能的衰减速率。
  • 分析失败类型,揭示模型在识别、空间、符号和语义上的脆弱点。

视觉语言模型(VLMs)在真实场景中常因输入图像质量低或不稳定而表现不佳。本文提出BRUCE(Benchmarking Robustness Under Corruption Escalation),一个针对科学视觉语言推理的多模态推理脆弱性评估框架。现有评估主要关注干净任务准确率,忽视推理稳定性在不同鲁棒性维度下的退化。BRUCE在多种图像扰动(如模糊、低对比度)下进行测试,并引入两个新指标:鲁棒性退化指数(RCI)和遍历-RCI(T-RCI),用于量化多模态推理性能随视觉退化严重程度提升时的衰减速率。我们在化学与数学推理任务上对多个数据集进行评估,按四大高阶推理领域分析退化引发的预测失败:依赖OCR的推理、空间推理、符号推理和语义错误,每类包含细粒度的特定退化失败子类型,实现可解释的故障分析。

原文摘要 · Abstract (English)

Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this paper, we aim at analyzing VLMs' robustness by applying perturbations and distortions to the input images, such as blur or low contrast. Toward this goal, we propose BRUCE (Benchmarking Robustness Under Corruption Escalation), a multimodal reasoning fragility framework for scientific vision-language reasoning. State-of-the-art evaluation frameworks/studies primarily focus on clean-task accuracy and rarely analyze how reasoning stability degrades across robustness dimensions. Besides varying over a wide-range of input perturbations, BRUCE employs two novel metrics -- Robustness Corruption Index (RCI) and Traversal-RCI (T-RCI) -- to quantify how rapidly multimodal reasoning performance deteriorates in VLMs as visual corruption severity increases under progressive perturbation scaling. We evaluate BRUCE across chemistry and mathematical reasoning tasks for multiple datasets, while analyzing corruption-induced prediction failures in terms of four high-level reasoning domains: OCR-dependent reasoning, spatial reasoning, symbolic reasoning, and semantic failures, with each containing fine-grained corruption specific failure subtypes, thereby enabling an interpretable failure analysis.

视觉语言模型鲁棒性评测科学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。