arXiv:2607.22035cs.LG2026-07

通过条件敏感性检测模型对版权内容的记忆,区分侵权与正常相似。

DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection

论文配图:DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection
图 1 · 摘自论文原文
  • 构建双分支框架,用反事实扰动衡量模型对版权内容的依赖程度。
  • 在多个模型上验证,能有效识别出目标特定记忆而非泛化偏差。
  • 适合评估生成模型版权风险,尤其对需合规的AI应用者有参考价值。

当前多数基础模型会复现或高度依赖受版权保护的训练内容,但仅靠输出相似性不足以判定侵权,因为相似结果也可能源于公共领域概念、常见风格或普通统计泛化。本文提出一种统一的后处理检测框架,将版权侵权证据视为反事实条件分布偏移:当目标内容被加入或移除训练过程时,若模型在对齐条件下行为发生显著变化,则该目标可疑。我们基于条件差分隐私形式化这一观点,引入双分支条件敏感性(DCS)——一种度量两个局部扰动模型状态间可观测差距的操作统计量。具体而言,该框架在部署模型周围构建学习分支与遗忘分支,通过影响函数分析将两者位移关联至不可观测的反事实重训练效应,并以反事实隐私预算代理、局部曲率、训练集规模和扰动步长为边界约束可观测敏感性。为区分目标特定记忆与通用微调不稳定性,进一步定义校准检测统计量,即从正交条件下测得的敏感性中减去基线值。该框架已应用于岭回归线性模型、条件扩散模型、自回归语言模型及多模态模型,展示了通过预测差异、图像嵌入发散、词元分布或熵变化、跨模态表示转移等不同方式实现同一原则的可行性。

原文摘要 · Abstract (English)

Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain concepts, common stylistic conventions, or ordinary statistical generalization. In this paper, we develops a unified post-hoc detection framework that treats copyright infringement evidence as a counterfactual conditional distribution shift: a protected target is suspicious when the model's behavior under aligned conditions would change measurably if that target were included in, or removed from, the training process. We formalize this view through conditional differential privacy and introduce Dual-Branch Conditional Sensitivity (DCS), an operational statistic that measures the observable gap between two locally perturbed model states. Specifically, the proposed DCS framework creates a learning branch and an unlearning branch around the deployed model, connects their displacement to the unavailable counterfactual retraining effect through influence-function analysis, and bounds the observable sensitivity by the counterfactual privacy-budget surrogate, local curvature, training-set scale, and perturbation step size. To distinguish target-specific memorization from generic fine-tuning instability, we further define a calibrated detection statistic that subtracts the sensitivity measured under orthogonal conditions. The DCS framework is instantiated for ridge-regularized linear regression, conditional diffusion models, autoregressive language models, and multimodal models. These instantiations show how the same principle can be evaluated through prediction gaps, image-embedding divergence, token-distribution or entropy shifts, and cross-modal representation changes.

版权检测生成模型敏感性分析反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。