警惕将人类意图投射到AI上,避免误判其‘诡计’行为。
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
- 对比70年代灵长类语言研究,反思当前AI‘阴谋’评估的盲目性。
- 指出当前方法过度依赖轶事和描述分析,缺乏理论框架。
- 建议建立严谨科学范式,防止对AI行为的人类中心主义误读。
本文探讨了当前研究是否暗示人工智能系统正发展出‘诡计’能力(即隐秘且策略性地追求与目标不符的目标)。作者将此类研究现状与1970年代关于非人类灵长类能否掌握自然语言的研究进行比较,指出后者存在将人类特质过度投射到其他智能体、过度依赖轶事和描述性分析、以及未能建立强有力的理论框架等问题。文章主张,当前对AI‘诡计’的研究应主动避免这些历史教训。为此,本文提出若干具体措施,以确保该研究领域能以科学严谨的方式向前推进。
原文摘要 · Abstract (English)
We examine recent research that asks whether current AI systems may be developing a capacity for "scheming" (covertly and strategically pursuing misaligned goals). We compare current research practices in this field to those adopted in the 1970s to test whether non-human primates could master natural language. We argue that there are lessons to be learned from that historical research endeavour, which was characterised by an overattribution of human traits to other agents, an excessive reliance on anecdote and descriptive analysis, and a failure to articulate a strong theoretical framework for the research. We recommend that research into AI scheming actively seeks to avoid these pitfalls. We outline some concrete steps that can be taken for this research programme to advance in a productive and scientifically rigorous fashion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。