让智能体用解释减少人类干预次数,提升协作效率。
"Trust me on this" Explaining Agent Behavior to a Human Terminator
- 通过可解释性机制说明行为理由,降低人为接管频率。
- 实验证明能显著减少干预次数,同时保持任务成功率。
- 适合自动驾驶、医疗等需人机协同的高风险场景。
在人机协作场景中,预训练智能体运行时,人类操作员可临时终止其操作并接管。这类场景常见于自动驾驶、工厂自动化和医疗领域。通常面临两难:若不允许接管,智能体可能采取次优甚至危险策略;若接管过多,人类对智能体失去信任,削弱其价值。本文形式化该设置,并提出一种可解释性方案,帮助优化人类干预次数,实现高效人机协同。
原文摘要 · Abstract (English)
Consider a setting where a pre-trained agent is operating in an environment and a human operator can decide to temporarily terminate its operation and take-over for some duration of time. These kind of scenarios are common in human-machine interactions, for example in autonomous driving, factory automation and healthcare. In these settings, we typically observe a trade-off between two extreme cases -- if no take-overs are allowed, then the agent might employ a sub-optimal, possibly dangerous policy. Alternatively, if there are too many take-overs, then the human has no confidence in the agent, greatly limiting its usefulness. In this paper, we formalize this setup and propose an explainability scheme to help optimize the number of human interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。