arXiv:2410.01927cs.CYcs.AI2024-10被引 14

如何让智能体AI的风险态度与人类用户匹配,避免失控或责任真空。

Risk Alignment in Agentic AI Systems

  • 提出风险对齐框架,使智能体决策符合用户风险偏好。
  • 指出风险态度错配可能导致系统失控或责任归属缺失。
  • 适合关注AI安全、伦理设计及自主系统监管的研究者。

具有自主行动能力的智能体AI正带来新的技术前沿,同时也引发安全性与对齐性新挑战。由于智能体的行为受其风险态度影响,风险对齐成为关键问题。若智能体风险偏好过激(源于对高风险用户的误校准或设计缺陷),可能造成严重威胁,并产生‘责任空白’——无人可为有害行为负责。本文探讨三个核心问题:智能体应遵循何种风险态度?如何使其与用户风险偏好对齐?是否应设定风险态度的边界?以及代他人做风险决策时的伦理考量。研究聚焦于这一领域的规范性与技术性议题。

原文摘要 · Abstract (English)

Agentic AIs $-$ AIs that are capable and permitted to undertake complex actions with little supervision $-$ mark a new frontier in AI capabilities and raise new questions about how to safely create and align such systems with users, developers, and society. Because agents' actions are influenced by their attitudes toward risk, one key aspect of alignment concerns the risk profiles of agentic AIs. Risk alignment will matter for user satisfaction and trust, but it will also have important ramifications for society more broadly, especially as agentic AIs become more autonomous and are allowed to control key aspects of our lives. AIs with reckless attitudes toward risk (either because they are calibrated to reckless human users or are poorly designed) may pose significant threats. They might also open 'responsibility gaps' in which there is no agent who can be held accountable for harmful actions. What risk attitudes should guide an agentic AI's decision-making? How might we design AI systems that are calibrated to the risk attitudes of their users? What guardrails, if any, should be placed on the range of permissible risk attitudes? What are the ethical considerations involved when designing systems that make risky decisions on behalf of others? We present three papers that bear on key normative and technical aspects of these questions.

智能体AI风险对齐人工智能伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。