让AI在保证安全的前提下,解释自身行为,提升人机协作信任度。
Safe Explicable Policy Search
- 结合约束策略优化与可解释策略搜索,生成安全且可解释的行为。
- 在Safety-Gym环境和真实机器人上验证,既保安全又提升可解释性。
- 适合需要人机协同的机器人应用,尤其关注安全与透明性的场景。
当用户与智能体交互时,会形成对智能体的预期。满足这些预期对成功的人机互动至关重要,但用户预期可能与智能体的实际行为不一致。为此,需在规划中引入两种独立决策模型以生成可解释行为。然而,现有方法缺乏对安全性的考量,尤其是在学习环境中。本文提出安全可解释策略搜索(SEPS),旨在通过学习生成既安全又可解释的行为,同时在训练过程和训练后均控制风险。将SEPS建模为约束优化问题,目标是在满足安全约束和基于模型的次优性标准前提下,最大化可解释性得分。该方法创新性融合约束策略优化与可解释策略搜索,首次实现连续状态与动作空间下安全可解释行为的生成,对机器人应用具有重要意义。在Safety-Gym环境及物理机器人实验中验证了其有效性与实际价值。
原文摘要 · Abstract (English)
When users work with AI agents, they form conscious or subconscious expectations of them. Meeting user expectations is crucial for such agents to engage in successful interactions and teaming. However, users may form expectations of an agent that differ from the agent's planned behaviors. These differences lead to the consideration of two separate decision models in the planning process to generate explicable behaviors. However, little has been done to incorporate safety considerations, especially in a learning setting. We present Safe Explicable Policy Search (SEPS), which aims to provide a learning approach to explicable behavior generation while minimizing the safety risk, both during and after learning. We formulate SEPS as a constrained optimization problem where the agent aims to maximize an explicability score subject to constraints on safety and a suboptimality criterion based on the agent's model. SEPS innovatively combines the capabilities of Constrained Policy Optimization and Explicable Policy Search to introduce the capability of generating safe explicable behaviors to domains with continuous state and action spaces, which is critical for robotic applications. We evaluate SEPS in safety-gym environments and with a physical robot experiment to show its efficacy and relevance in human-AI teaming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。