arXiv:2502.00375cs.CV2025-02被引 3

跨模态识别AI生成内容,无需重训练即可适配新模型。

Scalable Framework for Classifying AI-Generated Content Across Modalities

  • 结合感知哈希与相似性度量,实现多模态内容分类。
  • 在Defactify4数据集上准确区分人类与AI生成内容。
  • 框架可扩展,适合快速部署于新生成模型检测场景。

生成式AI技术的快速发展使得区分人类与AI生成内容、以及识别不同生成模型输出变得尤为重要。本文提出一种可扩展框架,融合感知哈希、相似性度量与伪标签机制,实现对文本与图像等多模态内容的高效分类。该方法可在不重新训练的情况下集成新生成模型,具备良好的适应性与鲁棒性。在Defactify4数据集上的全面评估显示,该框架在文本与图像分类任务中均表现优异,能高精度区分人类与AI生成内容,并有效识别不同生成方法。结果表明,该框架具有在真实场景中应对持续演进的生成式AI的潜力。源代码已公开于 https://github.com/ffyyytt/defactify4。

原文摘要 · Abstract (English)

The rapid growth of generative AI technologies has heightened the importance of effectively distinguishing between human and AI-generated content, as well as classifying outputs from diverse generative models. This paper presents a scalable framework that integrates perceptual hashing, similarity measurement, and pseudo-labeling to address these challenges. Our method enables the incorporation of new generative models without retraining, ensuring adaptability and robustness in dynamic scenarios. Comprehensive evaluations on the Defactify4 dataset demonstrate competitive performance in text and image classification tasks, achieving high accuracy across both distinguishing human and AI-generated content and classifying among generative methods. These results highlight the framework's potential for real-world applications as generative AI continues to evolve. Source codes are publicly available at https://github.com/ffyyytt/defactify4.

内容识别多模态生成检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。