用大模型推理文本训练轻量检测器,精准识别高保真伪造图像。
ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation

- 将大模型生成的推理文本转化为可迁移的语义特征
- 在三个数据集上超越现有方法,尤其对复杂伪造图像表现更优
- 适合需要高效部署且关注语义敏感性的安全检测场景
AI生成图像(AIGI)的兴起对数字真实性构成严峻挑战,亟需高效、可泛化的图像伪造检测系统。现有方法分非大模型与大模型两类:前者擅长低级伪影检测但缺乏语义理解,后者具备强语义推理与可解释性但计算开销大且对细微视觉伪影不敏感。本文研究大模型生成推理文本的内在价值,将其视为提升泛化性与语义敏感性的关键。据此提出ReAlign框架,通过对比学习将GRPO优化的大模型生成高质量推理文本蒸馏为轻量级AIGI检测器。ReAlign继承了推理文本的泛化能力与语义敏感性,同时保持高效轻量。采用图像-文本对齐与分类损失联合优化策略。在AIGCDetectBenchmark、AIGI-Holmes及新构建的UltraSynth-10k数据集上的实验表明,ReAlign在准确率与泛化性能上持续优于当前最优检测器,尤其在现代生成模型产生的复杂高保真伪造图像面前表现突出。
原文摘要 · Abstract (English)
The rise of AI-generated images (AIGIs) poses growing challenges for digital authenticity, prompting the need for efficient, generalizable image forgery detection systems. Existing methods, whether non-LLM-based or LLM-based, exhibit distinct advantages and limitations. While non-LLM-based models offer efficient low-level artifact detection, they often lack semantic understanding. Conversely, LLM-based methods provide strong semantic reasoning and explainability but are computationally intensive and less sensitive to subtle visual artifacts. Moreover, the true contribution of explanatory reasoning texts to forgery detection performance remains unclear. In this work, we investigate the intrinsic value and potential of LLM-generated reasoning texts, considering it a source of generalization and semantic-error sensitivity. Based on these findings, we propose ReAlign, a novel framework that distills high-quality reasoning texts generated by a GRPO-optimized LLM into a lightweight AIGI detector via contrastive learning. ReAlign effectively inherits the generalization ability and semantic sensitivity capability of reasoning textual representations, while remaining efficient and lightweight for deployment. Moreover, ReAlign adopts a tailored joint optimization strategy that integrates contrastive loss for image-text alignment and classification loss for accurate forgery discrimination. Experimental results on AIGCDetectBenchmark, AIGI-Holmes, and our newly constructed UltraSynth-10k demonstrate that ReAlign consistently outperforms existing state-of-the-art detectors in both accuracy and generalization, particularly when facing complex, high-fidelity forgeries from modern generative models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。