arXiv:2506.16407cs.CVcs.AI2025-06被引 4

首次系统评估多模态攻击对文字识别文档理解模型的破坏力

Robustness Evaluation of OCR-based Visual Document Understanding under Multi-Modal Adversarial Attacks

  • 构建统一框架生成六类基于梯度的布局攻击,涵盖框、像素、文本多粒度扰动
  • 线级攻击与复合扰动使模型性能下降超40%,且保持布局合理性(IoU≥0.6)
  • 适用于评估和提升文档理解模型在真实对抗环境下的鲁棒性

视觉文档理解(VDU)系统通过融合文本、版式和视觉信号,在信息抽取任务中表现优异。然而,其在真实对抗扰动下的鲁棒性仍缺乏充分研究。本文提出首个统一框架,用于生成与评估基于OCR的VDU模型面临的多模态对抗攻击。方法涵盖六种基于梯度的版式攻击场景,对文字框、像素及文本进行词级与行级扰动,并设置布局扰动预算约束(如IoU ≥ 0.6)以保证语义合理性。在四大数据集(FUNSD, CORD, SROIE, DocVQA)和六类模型上的实验表明,行级攻击与复合扰动(BBox + Pixel + Text)导致最严重性能下降。基于PGD的框扰动在所有模型中均优于随机平移基线。消融实验验证了布局预算、文本修改与对抗迁移性的关键影响。

原文摘要 · Abstract (English)

Visual Document Understanding (VDU) systems have achieved strong performance in information extraction by integrating textual, layout, and visual signals. However, their robustness under realistic adversarial perturbations remains insufficiently explored. We introduce the first unified framework for generating and evaluating multi-modal adversarial attacks on OCR-based VDU models. Our method covers six gradient-based layout attack scenarios, incorporating manipulations of OCR bounding boxes, pixels, and texts across both word and line granularities, with constraints on layout perturbation budget (e.g., IoU >= 0.6) to preserve plausibility. Experimental results across four datasets (FUNSD, CORD, SROIE, DocVQA) and six model families demonstrate that line-level attacks and compound perturbations (BBox + Pixel + Text) yield the most severe performance degradation. Projected Gradient Descent (PGD)-based BBox perturbations outperform random-shift baselines in all investigated models. Ablation studies further validate the impact of layout budget, text modification, and adversarial transferability.

文档理解对抗攻击OCR鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。