让AI学会合理反抗,提升人机协作安全与效率
Artificial Intelligent Disobedience: Rethinking the Agency of Our Artificial Teammates
- 提出AI自主性分级体系,定义智能违抗的边界
- 实证表明盲目服从的AI可能危及任务安全
- 适合研究人机协作与AI伦理的学者参考
近年来,人工智能在众多任务中已实现超人类表现。然而,大多数协作型AI系统仍过于顺从,严格遵循指令而无视潜在风险或失效。本文主张扩展AI队友的自主性,引入“智能违抗”机制,使其能在人机协同中主动干预不合理指令。论文构建了AI自主性层级量表,并通过典型案例说明在不同自主水平下违抗行为的表现形式。最后提出研究智能违抗需关注的初始边界与考量因素,强调将自主性作为合作场景下的独立研究方向。
原文摘要 · Abstract (English)
Artificial intelligence has made remarkable strides in recent years, achieving superhuman performance across a wide range of tasks. Yet despite these advances, most cooperative AI systems remain rigidly obedient, designed to follow human instructions without question and conform to user expectations, even when doing so may be counterproductive or unsafe. This paper argues for expanding the agency of AI teammates to include \textit{intelligent disobedience}, empowering them to make meaningful and autonomous contributions within human-AI teams. It introduces a scale of AI agency levels and uses representative examples to highlight the importance and growing necessity of treating AI autonomy as an independent research focus in cooperative settings. The paper then explores how intelligent disobedience manifests across different autonomy levels and concludes by proposing initial boundaries and considerations for studying disobedience as a core capability of artificial agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。