提出统一隐码恢复框架,实现自然图像伪造内容的还原与事实召回。
Beyond Detection: Multi-Scale Hidden-Code for Natural Image Deepfake Recovery and Factual Retrieval
- 通过多尺度向量量化编码语义与感知信息,生成紧凑隐码表示。
- 在ImageNet-S数据集上实现高精度的事实召回与图像重建。
- 兼容多种水印方案,适用于后处理与生成中水印场景。
近年来图像真实性研究主要聚焦于深度伪造检测与定位,对篡改内容的恢复与事实召回仍缺乏探索。本文提出一个统一的隐码恢复框架,支持从后处理与生成过程中水印两种范式中实现内容恢复与事实召回。方法将语义与感知信息编码为紧凑的隐码表示,通过多尺度向量量化进行优化,并利用条件Transformer模块增强上下文推理能力。为系统评估自然图像的恢复效果,构建了ImageNet-S基准,提供成对的图像-标签事实召回任务。在ImageNet-S上的大量实验表明,该方法在事实召回与图像重建方面表现优异,同时与多种水印流程完全兼容。该框架为超越检测与定位的通用图像恢复奠定了基础。
原文摘要 · Abstract (English)
Recent advances in image authenticity have primarily focused on deepfake detection and localization, leaving recovery of tampered contents for factual retrieval relatively underexplored. We propose a unified hidden-code recovery framework that enables both retrieval and restoration from post-hoc and in-generation watermarking paradigms. Our method encodes semantic and perceptual information into a compact hidden-code representation, refined through multi-scale vector quantization, and enhances contextual reasoning via conditional Transformer modules. To enable systematic evaluation for natural images, we construct ImageNet-S, a benchmark that provides paired image-label factual retrieval tasks. Extensive experiments on ImageNet-S demonstrate that our method exhibits promising retrieval and reconstruction performance while remaining fully compatible with diverse watermarking pipelines. This framework establishes a foundation for general-purpose image recovery beyond detection and localization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。