为社交平台生成内容设计可追溯的水印与多模态有害检测系统。
Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection

- 在图像生成时嵌入加密标识,实现内容源头追踪。
- 波段域扩频水印在模糊干扰下仍保持强鲁棒性。
- 结合图文检测模型,精准识别跨模态有害内容,适合监管与平台风控。
生成式AI的快速发展给内容审核与数字取证带来新挑战。例如,良性生成图像与有害或误导性文本结合,形成难以察觉的滥用行为,破坏传统审核机制并使溯源困难,因合成图像通常缺乏持久元数据或设备标识。本文提出一种基于隐写术的溯源框架,在图像生成时嵌入加密签名标识,并利用多模态有害内容检测作为溯源验证触发器。系统评估了五种水印方法在空间、频域和小波域的表现,同时集成基于CLIP的融合模型进行多模态有害内容检测。实验表明,小波域扩频水印在模糊干扰下具有强鲁棒性,多模态融合检测器达到0.99的AUC-ROC,支持可靠的跨模态溯源验证。该系统构建端到端取证流程,助力现代合成媒体环境下的责任追溯。代码已开源:https://github.com/bli1/steganography
原文摘要 · Abstract (English)
The rapid growth of generative AI has introduced new challenges in content moderation and digital forensics. In particular, benign AI-generated images can be paired with harmful or misleading text, creating difficult-to-detect misuse. This contextual misuse undermines the traditional moderation framework and complicates attribution, as synthetic images typically lack persistent metadata or device signatures. We introduce a steganography enabled attribution framework that embeds cryptographically signed identifiers into images at creation time and uses multimodal harmful content detection as a trigger for attribution verification. Our system evaluates five watermarking methods across spatial, frequency, and wavelet domains. It also integrates a CLIP-based fusion model for multimodal harmful-content detection. Experiments demonstrate that spread-spectrum watermarking, especially in the wavelet domain, provides strong robustness under blur distortions, and our multimodal fusion detector achieves an AUC-ROC of 0.99, enabling reliable cross-modal attribution verification. These components form an end-to-end forensic pipeline that enables reliable tracing of harmful deployments of AI-generated imagery, supporting accountability in modern synthetic media environments. Our code is available at GitHub: https://github.com/bli1/steganography
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。