AI安全研究需更强证据,避免误判模型风险。
Position: Anthropomorphic Misalignment Research Needs Stronger Evidence

- 提出证据等级框架与诊断清单,提升研究严谨性。
- 指出概念模糊、数据不稳等导致行为误读。
- 适合关注AI安全评估与政策制定的研究者。
我们主张,许多人格化错位研究(AMR)需更强证据支持,以确保其能为关键安全决策(如模型部署与监管)提供可靠基础。通过分析欺骗、涌现错位、逢迎等不同错位概念的失效模式,我们揭示了概念模糊、非鲁棒数据集、实验设计缺陷及不足因果干预如何导致对模型行为的过度解读。本文旨在提供证据考量指引,推动方法论改进。为此,我们提出证据等级框架与诊断检查表,建立共享标准,促进更有效的科学讨论,确保关于AI风险的主张建立在坚实的实证基础上。
原文摘要 · Abstract (English)
We argue that many Anthropomorphic Misalignment Research (AMR) studies need stronger evidence to ensure that they can provide a robust foundation for critical safety decisions, such as model deployment and regulation. By evaluating failure modes across different misalignment concepts, such as deception, emergent misalignment, and sycophancy, we show how conceptual ambiguity, non-robust datasets, experimental design, and insufficient causal interventions can lead to overinterpretation of model behaviors. This position paper aims to offer guidance on evidentiary considerations that can help improve methodological rigor in AMR. To achieve this, we provide a clear call to action through a proposed framework of evidence levels and a diagnostic checklist. These shared standards will enable more productive scientific discourse and ensure that claims about AI risks rest on solid empirical foundations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。