提升视觉退化科学图的推理能力,让AI更抗干扰。
Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams
- 多视角生成+一致性验证,自动修正模型错误
- 退化图上准确率从85.2%降至72.1%,暴露现有模型弱点
- 专为抗干扰设计的新数据集与评测指标,适合科研工具开发者
大型语言模型(LLMs)及其多模态变体(LVLMs)在科学和工程领域有巨大潜力,尤其在处理科学图表等视觉信息方面。然而,其实际应用受限于对噪声、模糊和遮挡等常见视觉退化现象的脆弱性,这些在真实科学文档中普遍存在。现有评估基准大多忽略此问题,导致LVLM在视觉退化的科学图表上的鲁棒推理能力未被充分研究。为此,我们提出鲁棒图表推理(RDR)框架,旨在增强并严格评估LVLM在这些条件下的性能。核心是自适应多视图与一致性验证(AMCV)机制:生成多个退化版本图表,进行并行推理,并通过基于一致性的自校正循环优化结果。我们还提出了两个新指标——扰动鲁棒得分(PRS)和视觉退化一致性(VDC),用于量化鲁棒性。此外,构建了首个大规模科学图表问答数据集SciDiagram-Robust,通过程序化生成多种视觉退化进行增强。大量实验表明,即使是最先进的闭源模型如GPT-4V,在面对退化输入时性能显著下降(干净准确率85.2%对比PRS 72.1%)。
原文摘要 · Abstract (English)
Large Language Models (LLMs) and their multimodal variants (LVLMs) hold immense promise for scientific and engineering applications, particularly in processing visual information like scientific diagrams. However, their practical deployment is hindered by a critical lack of robustness to common visual perturbations such as noise, blur, and occlusions, which are prevalent in real-world scientific documents. Existing evaluation benchmarks largely overlook this challenge, leaving the robust reasoning capabilities of LVLMs on visually degraded scientific diagrams underexplored. To address this, we introduce the Robust Diagram Reasoning (RDR) framework, a novel approach designed to enhance and rigorously evaluate LVLMs' performance under such conditions. At its core, RDR employs an Adaptive Multi-View & Consistency Verification (AMCV) mechanism, which involves generating multiple perturbed versions of a diagram, performing parallel inference, and then applying a consistency-based self-correction loop. We also propose two new metrics, Perturbation Robustness Score (PRS) and Visual Degradation Consistency (VDC), to quantify robustness. Furthermore, we construct SciDiagram-Robust, the first large-scale scientific diagram question-answering dataset specifically augmented with diverse, programmatically generated visual perturbations. Our extensive experiments demonstrate that even state-of-the-art closed-source LVLMs like GPT-4V exhibit significant performance degradation when faced with perturbed inputs (Clean Accuracy 85.2% vs. PRS 72.1%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。