arXiv:2411.10461cs.HCcs.AI2024-11NeurIPS被引 14

用行为模型操控AI解释,能轻易引导人做特定决策。

Utilizing Human Behavior Modeling to Manipulate Explanations in AI-Assisted Decision Making: The Good, the Bad, and the Scary

  • 构建人类融合AI建议与解释的行为量化模型。
  • 操纵解释可有效引导决策,成功率高且人难察觉。
  • 适用于提升效率或隐蔽干预场景,需警惕滥用风险。

近期人工智能模型的发展推动了AI决策辅助在人类决策过程中的广泛应用。为充分发挥人机协同潜力,研究者已建立计算模型以理解人类如何整合AI建议并据此做出最终决策,进而优化人机协作表现。同时,由于AI模型的“黑箱”特性,向人类决策者提供解释以帮助其更合理地依赖推荐已成为常见做法。本文探讨能否定量建模人类如何将AI建议与解释共同融入决策流程,并利用该模型生成可操控的解释,从而引导个体达成特定目标。我们在多个任务中开展大规模人类实验,结果表明:无论意图是恶意还是善意,被操纵的解释都能轻易影响人类行为并实现预期结果;且参与者通常无法察觉解释中的异常,即便其决策已被改变。

原文摘要 · Abstract (English)

Recent advances in AI models have increased the integration of AI-based decision aids into the human decision making process. To fully unlock the potential of AI-assisted decision making, researchers have computationally modeled how humans incorporate AI recommendations into their final decisions, and utilized these models to improve human-AI team performance. Meanwhile, due to the ``black-box'' nature of AI models, providing AI explanations to human decision makers to help them rely on AI recommendations more appropriately has become a common practice. In this paper, we explore whether we can quantitatively model how humans integrate both AI recommendations and explanations into their decision process, and whether this quantitative understanding of human behavior from the learned model can be utilized to manipulate AI explanations, thereby nudging individuals towards making targeted decisions. Our extensive human experiments across various tasks demonstrate that human behavior can be easily influenced by these manipulated explanations towards targeted outcomes, regardless of the intent being adversarial or benign. Furthermore, individuals often fail to detect any anomalies in these explanations, despite their decisions being affected by them.

人机协同可解释AI行为建模决策操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。