测试视觉干扰下视觉语言模型的OCR推理鲁棒性,发现高准确率不等于强抗扰能力。
How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbations

- 构建OCR-Robust基准,涵盖文档、图表、表格等812个样本
- 5种干扰类型在3个强度下测试,发现图表类任务更易受破坏
- 模型在结构敏感任务中退化严重,清洁准确率高未必抗扰性强
视觉语言模型(VLMs)在基于OCR的基准上表现优异,日益关注文本密集型理解,但其在可控视觉退化下的鲁棒性仍不清楚。这一缺口对OCR推理至关重要,因为视觉污染可能导致OCR错误和结构扭曲,从而引入推理不确定性。为此,我们提出OCR-Robust基准,用于评估视觉扰动下的OCR推理鲁棒性。该基准包含812个样本,分为两个互补子集:OCR1.0覆盖文档、场景文字、收据、手写体及数学内容;OCR2.0聚焦图表、几何图示和表格。通过初步研究18种候选扰动,筛选出5种代表性类型,每种设3个严重等级。采用清洁准确率、相对扰动保留率(RCR)、最差情况保留率(WCR)和综合扰动鲁棒性指数(CRI)评估18个模型,涵盖专有系统、开源VLMs及OCR+LLM流水线。结果表明,更高清洁准确率并不意味着更强鲁棒性,且对结构敏感的任务在最坏情况下退化显著,图表与表格比文档类输入更脆弱。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have achieved strong performance on OCR-based benchmarks and increasingly focused on text-rich understanding, but their robustness under controlled visual degradation remains insufficiently understood. This gap is critical for OCR reasoning, where visual corruption can induce OCR errors and structural distortions, thereby introducing uncertainty into the reasoning task. To systematically study this problem, we introduce OCR-Robust, a benchmark designed for evaluating OCR reasoning robustness under visual perturbations. It contains 812 samples across two complementary subsets: OCR1.0, covering documents, scene text, receipts, handwriting, and mathematical content, and OCR2.0, focusing on charts, geometry diagrams, and tables. To enable efficient yet informative evaluation, we conduct a pilot study over 18 candidate perturbations and select 5 representative types at 3 severity levels each based on their impact and cross-model discriminability. We evaluate robustness using clean accuracy, Relative Corruption Retention (RCR), Worst-Case Retention (WCR), and a composite Corruption Robustness Index (CRI), and benchmark 18 models spanning proprietary systems, open-source VLMs, and OCR+LLM pipelines. Our results show that higher clean accuracy does not necessarily imply stronger robustness, and that models can suffer pronounced degradation in the worst case on OCR tasks that are sensitive to structure, and charts and tables are substantially more fragile than document-like inputs under perturbation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。