arXiv:2512.21482cs.AI2025-12被引 2

用视觉逻辑协同推理检测图文伪造,提升准确率与可解释性。

LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis

  • 构建跨模态链式思维机制,交替验证图像线索与文本逻辑。
  • 在T-IC13零样本测试中F1得分领先专用模型41.4%。
  • 适合需要高精度图文真伪分析的AI安全与内容审核场景。

由AIGC快速发展驱动的复杂文本中心伪造内容,严重威胁社会安全与信息真实性。现有方法多依赖粗粒度视觉分析,缺乏深度推理能力,且将检测、定位与解释视为独立任务,忽视其内在关联。为此,我们提出LogicLens,一个统一的视觉-文本协同推理框架,将三者整合为联合任务。其核心是新颖的跨线索链式思维(CCT)机制,通过迭代交叉验证视觉线索与文本逻辑实现深层推理。为确保各任务间稳健对齐,引入基于GRPO优化的加权多任务奖励函数。同时,设计了分层迭代的多智能体系统PR²(Perceiver, Reasoner, Reviewer),生成高质量认知对齐标注;并构建了包含5,397张图像的RealText数据集,涵盖文本解释、像素级分割及真实性标签。大量实验表明,LogicLens在多个基准上表现卓越:在T-IC13的零样本评估中,宏平均F1分数超越专用框架41.4%,优于GPT-4o 23.4%;在挑战性密集文本的T-SROIE数据集上,于mF1、CSS及宏平均F1指标上显著领先其他多模态大模型方法。相关数据集、模型与代码将公开发布。

原文摘要 · Abstract (English)

Sophisticated text-centric forgeries, fueled by rapid AIGC advancements, pose a significant threat to societal security and information authenticity. Current methods for text-centric forgery analysis are often limited to coarse-grained visual analysis and lack the capacity for sophisticated reasoning. Moreover, they typically treat detection, grounding, and explanation as discrete sub-tasks, overlooking their intrinsic relationships for holistic performance enhancement. To address these challenges, we introduce LogicLens, a unified framework for Visual-Textual Co-reasoning that reformulates these objectives into a joint task. The deep reasoning of LogicLens is powered by our novel Cross-Cues-aware Chain of Thought (CCT) mechanism, which iteratively cross-validates visual cues against textual logic. To ensure robust alignment across all tasks, we further propose a weighted multi-task reward function for GRPO-based optimization. Complementing this framework, we first designed the PR$^2$ (Perceiver, Reasoner, Reviewer) pipeline, a hierarchical and iterative multi-agent system that generates high-quality, cognitively-aligned annotations. Then, we constructed RealText, a diverse dataset comprising 5,397 images with fine-grained annotations, including textual explanations, pixel-level segmentation, and authenticity labels for model training. Extensive experiments demonstrate the superiority of LogicLens across multiple benchmarks. In a zero-shot evaluation on T-IC13, it surpasses the specialized framework by 41.4% and GPT-4o by 23.4% in macro-average F1 score. Moreover, on the challenging dense-text T-SROIE dataset, it establishes a significant lead over other MLLM-based methods in mF1, CSS, and the macro-average F1. Our dataset, model, and code will be made publicly available.

伪造检测多模态推理视觉语言可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。