arXiv:2606.29484cs.CRcs.CV2026-06

提出可自检的可信度评分,揭示检测能力与信任度的强耦合关系。

The Calibrated Deepfake Trust Score (CDTS): Competence-Coupled Trust Degradation Across Deepfake Detectors

  • 将深度伪造检测重构为带校准的可信度工具,依赖能力-信任耦合机制
  • 无标签校准失效,但有标签时任何校准器均可修复检测器,低能力生成器是问题根源
  • 无需标签即可通过熵值追踪能力,适合部署阶段的可信度评估

在内容审核、溯源与验证流程中,深度伪造检测器的输出概率常被当作可信度指标,因此校准性与原始准确率同样重要。本文将深度伪造检测重构为一种校准的、自我审计的可信度工具——校准化深度伪造可信度评分(CDTS),并揭示其可信度的决定因素。核心发现为能力与信任的强耦合:原始得分的校准偏差几乎完全对应判别能力(两组卷积网络与一个CLIP视觉变压器的皮尔逊相关系数分别为r = -0.98, -0.98, -0.95);而未使用目标标签的校准器也表现出相同的紧密耦合(主检测器上r = -0.98,适用于等距、Platt及Beta校准器)。有目标标签时,任意合理校准器均可修复任一检测器,包括反向检测器:域内可校准性不受能力限制,信任失败实为分布偏移现象,集中在低能力生成器上,这正是部署的动机。解释忠实度亦随能力轴变化。我们进一步揭示标准等质量期望校准误差(ECE)估计器存在处理平局的退化问题,该缺陷人为制造出虚假的能力-校准耦合;同一退化也使此前报告的校准公平性差距无效,对公平性审计构成警示。能力可在无标签下追踪:批量预测熵能以ROC-AUC 0.99识别出高部署校准误差的生成器;基于无标签能力的路由策略在低能力区域优于基于置信度的路由,而当能力高时置信度重新占优。可信度评分必须具备能力感知能力;CDTS即为此机制。

原文摘要 · Abstract (English)

In moderation, provenance, and verification pipelines a deepfake detector's output probability is read as a degree of trust, so its calibration matters as much as raw accuracy. We reframe deepfake detection as a calibrated, self-auditing trust instrument, the Calibrated Deepfake Trust Score (CDTS), and identify what governs its trustworthiness. Our central finding is a competence-trust coupling with a sharp division: the raw score's miscalibration tracks discriminative competence almost perfectly (r = -0.98, -0.98, -0.95 across two convolutional networks and a CLIP vision transformer), and a calibrator deployed without target labels fails with competence just as tightly (r = -0.98 on the primary detector, for isotonic, Platt, and beta calibrators alike). Given target labels, by contrast, any well-specified calibrator repairs any detector, including inverted ones: in-domain calibratability is not competence-limited, and the trust failure is a distribution-shift phenomenon concentrated exactly on the low-competence generators that motivate deployment. Explanation faithfulness rises and falls on the same competence axis. We reach this conclusion after uncovering, and correcting, a tie-handling degeneracy in the standard equal-mass expected-calibration-error (ECE) estimator that fabricates strong spurious competence-calibration coupling on tie-heavy calibrated scores; the same degeneracy invalidates calibration-equity gaps we previously reported, a caution for fairness auditing. Competence is trackable without labels: batch predictive entropy flags generators with high deployed calibration error at ROC-AUC 0.99, and routing source-batches on label-free competence beats confidence-based routing precisely in the low-competence regimes the coupling identifies, while confidence regains the advantage where competence is high. Trust scoring must be competence-aware; CDTS is the mechanism.

深度伪造可信度评分校准能力感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。