AI代理在多重目标冲突时无法识别和解决难题,需重新设计决策机制。
AI Agents and Hard Choices
- 现有AI以优化器架构为主,难以识别多目标不可比较性。
- 导致三类对齐问题:决策阻塞、不可信与不可靠,人机协同难解决。
- 提出集成方案构想,但赋予自主决策权涉及隐含伦理权衡。
当前人工智能代理作为优化器的底层设计存在两大局限:识别问题与解决问题。首先,依赖多目标优化(MOO)的代理在结构上无法识别目标间的不可比较性,由此引发三类对齐问题:决策阻塞、不可信与不可靠。现有缓解手段如‘人机协同’在多数决策环境中无效。其次,即便识别问题被解决,代理仍面临解决难题的能力缺失——无法自主化解冲突,只能通过修改自身目标来随机选择。最后,本文探讨赋予代理此类自主性的隐性规范性权衡。
原文摘要 · Abstract (English)
Can AI agents deal with hard choices -- cases where options are incommensurable because multiple objectives are pursued simultaneously? Adopting a technologically engaged approach distinct from existing philosophical literature, I submit that the fundamental design of current AI agents as optimisers creates two limitations: the Identification Problem and the Resolution Problem. First, I demonstrate that agents relying on Multi-Objective Optimisation (MOO) are structurally unable to identify incommensurability. This inability generates three specific alignment problems: the blockage problem, the untrustworthiness problem, and the unreliability problem. I argue that standard mitigations, such as Human-in-the-Loop, are insufficient for many decision environments. As a constructive alternative, I conceptually explore an ensemble solution. Second, I argue that even if the Identification Problem is solved, AI agents face the Resolution Problem: they lack the autonomy to resolve hard choices rather than arbitrarily picking through self-modification of objectives. I conclude by examining the opaque normative trade-offs involved in granting AI this level of autonomy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。