用自举方法训练可信赖的AI裁判,让假视频检测理由更真实可信。
Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
- 通过自举生成-评估循环,将人类反馈转化为结构化推理监督
- 在新基准上达到96.2%准确率,推理相关性与人工评分高度一致
- 适合需要可解释、可信假视频检测的开发者与研究者
深度伪造检测模型常生成自然语言解释,但其推理往往缺乏视觉证据支撑,影响可靠性。现有评估仅关注分类准确率,忽略推理真实性。本文提出DeepfakeJudge框架,包含一个涵盖最新生成与编辑伪造的分布外基准、带有视觉推理标注的人工标注子集,以及一套无需显式真值推理的评估模型。该裁判通过自举的生成-评估流程优化,实现点对点和成对评估,可扩展人类反馈以形成结构化推理监督。在所提出的元评估基准上,我们的推理自举模型达到96.2%准确率,优于30倍大的基线模型。推理裁判与人工评分高度相关,在人工标注子集上达成98.9%的成对一致性。用户研究表明,70%参与者更偏好本框架生成的推理结果,因其更忠实、有据且有用。所有数据集、模型与代码已开源。
原文摘要 · Abstract (English)
Deepfake detection models often generate natural-language explanations, yet their reasoning is frequently ungrounded in visual evidence, limiting reliability. Existing evaluations measure classification accuracy but overlook reasoning fidelity. We propose DeepfakeJudge, a framework for scalable reasoning supervision and evaluation, that integrates an out-of-distribution benchmark containing recent generative and editing forgeries, a human-annotated subset with visual reasoning labels, and a suite of evaluation models, that specialize in evaluating reasoning rationales without the need for explicit ground truth reasoning rationales. The Judge is optimized through a bootstrapped generator-evaluator process that scales human feedback into structured reasoning supervision and supports both pointwise and pairwise evaluation. On the proposed meta-evaluation benchmark, our reasoning-bootstrapped model achieves an accuracy of 96.2\%, outperforming \texttt{30x} larger baselines. The reasoning judge attains very high correlation with human ratings and 98.9\% percent pairwise agreement on the human-annotated meta-evaluation subset. These results establish reasoning fidelity as a quantifiable dimension of deepfake detection and demonstrate scalable supervision for interpretable deepfake reasoning. Our user study shows that participants preferred the reasonings generated by our framework 70\% of the time, in terms of faithfulness, groundedness, and usefulness, compared to those produced by other models and datasets. All of our datasets, models, and codebase are \href{https://github.com/KjAeRsTuIsK/DeepfakeJudge}{open-sourced}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。