用逻辑推理检测文档伪造,让AI能像人一样找证据
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning

- 将图像与文本联合推理,自动发现文字篡改痕迹
- 在多个数据集上比现有方法提升超40%的识别准确率
- 适合需要高可信度文档安全验证的场景
生成式AI的发展使得以文本为中心的图像伪造日益逼真,严重威胁文档安全。现有检测方法主要依赖视觉线索,缺乏基于证据的逻辑推理能力,且检测、定位与解释常被割裂处理,影响结果可靠性与可解释性。为此,我们提出DocShield,首个将文本主导型伪造分析建模为视觉-逻辑协同推理问题的统一框架。其核心是新颖的跨线索链式思维(CCT)机制,通过迭代交叉验证视觉异常与文本语义,生成一致且有依据的鉴定结果。我们还引入基于GRPO优化的加权多任务奖励,对齐推理结构、空间证据与真实性判断。同时构建了多语言真实文档类文本图像数据集RealText-V1,包含像素级篡改掩码与专家级文本解释。大量实验表明,DocShield显著优于现有方法,在T-IC13上相比专用框架提升宏平均F1达41.4%,相比GPT-4o提升23.4%,在更具挑战性的T-SROIE基准上同样表现优异。相关数据集、模型与代码将公开发布。
原文摘要 · Abstract (English)
The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document safety. Existing forensic methods mainly rely on visual cues and lack evidence-based reasoning to reveal subtle text manipulations. Detection, localization, and explanation are often treated as isolated tasks, limiting reliability and interpretability. To tackle these challenges, we propose DocShield, the first unified framework formulating text-centric forgery analysis as a visual-logical co-reasoning problem. At its core, a novel Cross-Cues-aware Chain of Thought (CCT) mechanism enables implicit agentic reasoning, iteratively cross-validating visual anomalies with textual semantics to produce consistent, evidence-grounded forensic analysis. We further introduce a Weighted Multi-Task Reward for GRPO-based optimization, aligning reasoning structure, spatial evidence, and authenticity prediction. Complementing the framework, we construct RealText-V1, a multilingual dataset of document-like text images with pixel-level manipulation masks and expert-level textual explanations. Extensive experiments show DocShield significantly outperforms existing methods, improving macro-average F1 by 41.4% over specialized frameworks and 23.4% over GPT-4o on T-IC13, with consistent gains on the challenging T-SROIE benchmark. Our dataset, model, and code will be publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。