伪造的真假共存:人工与AI生成内容可同时通过双重认证
Authenticated Contradictions from Desynchronized Provenance and Watermarking

- 通过删减C2PA元数据字段,制造同时通过双验证的假证据
- 3500张图像测试中,新协议实现100%准确率识别冲突状态
- 适用于内容安全、数字版权领域研究者及平台审核人员
加密溯源标准如C2PA与不可见水印被视为内容认证的互补防御手段,但两者在技术上相互独立,互不依赖。本文首次形式化并实证揭示了‘完整性冲突’(Integrity Clash)现象:一个数字资产同时具备符合加密规则的C2PA声明(声称人为创作),其像素又携带识别为AI生成的水印,且两个信号在孤立验证下均能通过。我们构建了基于标准编辑流程的元数据清洗工作流,仅需语义上省略一个允许的断言字段即可生成此类被认证的伪造品,无需任何密码学突破。为填补该漏洞,提出跨层审计协议,联合评估溯源元数据与水印检测状态,在涵盖四个冲突矩阵状态及三种现实扰动条件的3,500张测试图像上实现100%分类准确率。结果表明,该验证层间的差距既无必要也极易修补。
原文摘要 · Abstract (English)
Cryptographic provenance standards such as C2PA and invisible watermarking are positioned as complementary defenses for content authentication, yet the two verification layers are technically independent: neither conditions on the output of the other. This work formalizes and empirically demonstrates the $\textit{Integrity Clash}$, a condition in which a digital asset carries a cryptographically valid C2PA manifest asserting human authorship while its pixels simultaneously carry a watermark identifying it as AI-generated, with both signals passing their respective verification checks in isolation. We construct metadata washing workflows that produce these authenticated fakes through standard editing pipelines, requiring no cryptographic compromise, only the semantic omission of a single assertion field permitted by the current C2PA specification. To close this gap, we propose a cross-layer audit protocol that jointly evaluates provenance metadata and watermark detection status, achieving 100% classification accuracy across 3,500 test images spanning four conflict-matrix states and three realistic perturbation conditions. Our results demonstrate that the gap between these verification layers is unnecessary and technically straightforward to close.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。