arXiv:2608.20122cs.CV2026-08

提升模型对对抗性文字的识别能力,让机器看得更准。

ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation

论文配图:ArmorOCR: Grounded Adversarial Visual Perception via Observation-Transferred Self-Distillation
图 1 · 摘自论文原文
  • 用自我蒸馏和策略优化,从人类可读的对抗文本中学习
  • 在390张图上实现高精度定位与识别,覆盖13种攻击类型
  • 适合安全敏感场景下的视觉感知系统研发者

大型多模态模型虽具备强大OCR能力,但对人类可读、模型难辨的对抗性视觉文本仍易受干扰。现有OCR基准多聚焦自然或文档文本,缺乏大规模、多任务且区域感知的对抗性评测。本文将对抗性OCR定义为“有根据的OCR感知”任务,提出首个面向该任务的基准数据集AdvSpot,包含390张带区域标注的图像,涵盖5大类、13种细粒度对抗性文本类型。为此,我们设计两阶段训练框架ArmorOCR:第一阶段通过有策略自我蒸馏(OPSD)从变换后的观察中获取缺失的对抗性感知;第二阶段利用组相对策略优化(GRPO)结合任务条件奖励,对定位、识别、完整检测及视觉问答进行精细化优化。在AdvSpot及其他对抗与通用OCR基准上的实验表明,ArmorOCR在保持良好通用OCR性能的同时,显著提升了对抗性文本的感知能力。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging for models to localize and recognize. Existing OCR benchmarks mainly focus on natural or document-style text, while adversarial OCR evaluations remain limited in scale, task coverage, or region-aware evaluation. In this paper, we formulate adversarial OCR as a \textbf{grounded OCR perception} task and introduce \textbf{AdvSpot}, the first benchmark for grounded adversarial OCR evaluation. AdvSpot comprises 390 images with region-level annotations, spanning 5 primary categories and 13 fine-grained adversarial OCR types. To address this challenge, we propose \textbf{ArmorOCR}, a two-stage training framework for robust adversarial OCR perception. ArmorOCR first acquires missing adversarial OCR perception from privileged transformed observations through On-Policy Self-Distillation (OPSD), and then refines grounded OCR perception through Group Relative Policy Optimization (GRPO) with task-conditioned rewards for localization, recognition, full spotting, and visual question answering (VQA). Experiments on our AdvSpot, other adversarial OCR benchmarks, and general OCR benchmarks demonstrate that ArmorOCR consistently improves adversarial OCR perception while preserving competitive general OCR capability.

OCR对抗样本多模态视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。