arXiv:2602.21716cs.CV2026-02被引 1

提升AI生成图像检测能力,解决特征融合中的注意力稀释问题。

TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image Detection

  • 设计轻量级融合适配器,利用最优传输机制融合纹理与语义特征。
  • 在多个主流模型上实现最高6%的检测准确率提升。
  • 适合关注AI内容安全、多模态模型优化的研究者使用。

AI生成图像(AIGI)技术快速发展,生成效果高度逼真,威胁公共信息真实性和安全性。近期研究证明,在多模态大语言模型(MLLMs)中结合纹理级伪影特征与语义特征,可增强AIGI检测能力。然而,我们初步分析发现,伪影特征具有高内部相似性,经softmax操作后导致注意力图近乎均匀,引发注意力稀释,阻碍语义与伪影特征的有效融合。为此,我们提出轻量级融合适配器TranX-Adapter,包含任务感知的最优传输融合模块,以伪影与语义预测概率间的Jensen-Shannon散度作为代价矩阵,实现伪影信息向语义特征的迁移;以及X-Fusion模块,通过交叉注意力机制实现语义信息向伪影特征的反向传递。在多个先进MLLM上的标准AIGI检测基准测试中,TranX-Adapter带来一致且显著的性能提升,最高达+6%准确率。

原文摘要 · Abstract (English)

Rapid advances in AI-generated image (AIGI) technology enable highly realistic synthesis, threatening public information integrity and security. Recent studies have demonstrated that incorporating texture-level artifact features alongside semantic features into multimodal large language models (MLLMs) can enhance their AIGI detection capability. However, our preliminary analyses reveal that artifact features exhibit high intra-feature similarity, leading to an almost uniform attention map after the softmax operation. This phenomenon causes attention dilution, thereby hindering effective fusion between semantic and artifact features. To overcome this limitation, we propose a lightweight fusion adapter, TranX-Adapter, which integrates a Task-aware Optimal-Transport Fusion that leverages the Jensen-Shannon divergence between artifact and semantic prediction probabilities as a cost matrix to transfer artifact information into semantic features, and an X-Fusion that employs cross-attention to transfer semantic information into artifact features. Experiments on standard AIGI detection benchmarks upon several advanced MLLMs, show that our TranX-Adapter brings consistent and significant improvements (up to +6% accuracy).

AI检测多模态特征融合生成内容

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。