提出多值多智能体责任归属模型,支持责任预判与价值对齐策略选择。
Responsibility in a Multi-Value Strategic Setting
- 构建多智能体多值场景下的责任归因框架
- 非占优后悔最小化策略可稳定降低预期责任度
- 适用于需伦理对齐的AI系统设计与决策分析
责任是多智能体系统及安全、可靠、伦理化AI的核心概念。然而,以往研究大多仅关注单一结果的责任归属。本文提出一种在多智能体、多值情境下的责任归因模型,并扩展该模型以实现责任预判,展示责任考量如何帮助智能体选择与其价值观一致的策略。特别地,我们证明非占优后悔最小化策略能稳定地最小化智能体的预期责任度。
原文摘要 · Abstract (English)
Responsibility is a key notion in multi-agent systems and in creating safe, reliable and ethical AI. However, most previous work on responsibility has only considered responsibility for single outcomes. In this paper we present a model for responsibility attribution in a multi-agent, multi-value setting. We also expand our model to cover responsibility anticipation, demonstrating how considerations of responsibility can help an agent to select strategies that are in line with its values. In particular we show that non-dominated regret-minimising strategies reliably minimise an agent's expected degree of responsibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。