arXiv:2512.11899cs.CV2025-12中稿 · ECCV被引 1

提出新评测框架,测试大模型读字与忽略干扰字的能力

Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models

  • 设计联合任务RIO-VQA,判断何时读字何时忽略干扰
  • 构建同场景反事实基准RIO-Bench,精准评估模型行为
  • 发现抗攻击与读字能力存在权衡,提供兼顾两者的防御基线

大型视觉语言模型(LVLMs)易受排版攻击,即图像中插入的误导性文本可覆盖视觉理解。然而现有评估与防御多聚焦于物体识别,忽视了文本阅读能力。这在现实中不可接受:真实场景常需同时识别物体与阅读场景文字(如识别行人时读取交通标志)。为此,我们提出新任务Read-or-Ignore VQA(RIO-VQA),要求模型根据上下文判断何时读取场景文字、何时忽略插入的干扰文本。为评估该能力,我们构建了RIO-Bench——一个同场景反事实基准,保持场景不变,仅改变问题意图(物体/文字)和文本状态(正常/攻击),从而减少混杂因素,实现模型行为的直接比较。实验揭示:现有以物体为中心的防御方法虽能通过抑制文本敏感性实现鲁棒性,却牺牲了文本阅读性能(即‘忽略’文本)。为此,我们提出数据驱动的防御基线,在RIO-Bench上同时提升鲁棒性与读字能力,补足此前仅忽略文本的基线。本工作揭示当前以物体为中心的鲁棒性研究与真实多模态需求间的根本错配,为构建可靠的大模型提供原则性路径。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) are vulnerable to typographic attacks, where misleading text inserted into an image can override visual understanding. However, existing evaluation protocols and defenses are largely focused on object recognition and do not consider text-reading capability. This is a critical oversight: real-world scenarios often require both recognizing objects and reading scene text (e.g., recognizing pedestrians while reading traffic signs), where simply ignoring all text for robustness is unacceptable in practice. To address this gap, we introduce a novel task, Read-or-Ignore VQA (RIO-VQA), which jointly evaluates both requirements: models must decide, from context, when to read scene text and when to ignore inserted distractor text. To evaluate this capability, we present RIO-Bench, a same-scene counterfactual benchmark that holds the scene fixed while varying only question intent (object vs. text) and text condition (clean vs. attack), enabling direct comparisons of model behaviors with reduced confounding factors. Using RIO-Bench, we highlight a trade-off: representative defenses developed in object-centric settings can achieve robustness by suppressing text sensitivity, at the cost of text-reading performance (i.e., "ignoring" text). Motivated by this trade-off, we provide a data-driven defense baseline that improves both requirements on RIO-Bench, complementing prior text-ignoring baselines. Overall, this work highlights a fundamental misalignment between the current object-centric robustness scope and real-world multimodal requirements, providing a principled path toward reliable LVLMs.

视觉语言模型对抗攻击文本识别评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。