现有AI欺骗检测器缺乏可靠评估数据,难以验证其有效性。
Difficulties with Evaluating a Deception Detector for AIs
- 提出评估欺骗检测器需有明确真假案例,但当前缺乏此类标注数据。
- 分析表明现有研究无法提供足够可信的欺骗/诚实样本用于测试。
- 适合关注AI安全与评估方法论的研究者阅读。
构建可靠的AI欺骗检测工具——能够预测AI系统是否在策略性地欺骗,而无需行为证据——对缓解高级AI系统的风险具有重要意义。然而,评估所提检测器的可靠性与有效性,需要我们能确信标注为欺骗或诚实的样本。本文认为,目前我们尚不具备必要的此类样本,并进一步识别出收集这些样本时存在的若干具体障碍。通过概念论证、对现有实证研究的分析以及新提出的典型案例分析,本文提供了相关证据。同时讨论了几种潜在的实证替代方案,认为尽管这些方法有一定价值,但单独使用仍显不足。因此,欺骗检测的进展需进一步深入思考这些问题。
原文摘要 · Abstract (English)
Building reliable deception detectors for AI systems -- methods that could predict when an AI system is being strategically deceptive without necessarily requiring behavioural evidence -- would be valuable in mitigating risks from advanced AI systems. But evaluating the reliability and efficacy of a proposed deception detector requires examples that we can confidently label as either deceptive or honest. We argue that we currently lack the necessary examples and further identify several concrete obstacles in collecting them. We provide evidence from conceptual arguments, analysis of existing empirical works, and analysis of novel illustrative case studies. We also discuss the potential of several proposed empirical workarounds to these problems and argue that while they seem valuable, they also seem insufficient alone. Progress on deception detection likely requires further consideration of these problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。