让机器人决策过程可解释,还能评估解释的可信度。
BaTCAVe: Trustworthy Explanations for Robot Behaviors
- 用人类能懂的概念匹配神经网络激活,生成解释。
- 解释附带置信度评分,可量化其可靠性。
- 适合机器人工程师和监管者做故障诊断。
黑箱神经网络是现代机器人不可或缺的部分。然而,在真实场景中部署此类高风险系统时,若工程师和立法机构无法理解其决策过程,将带来重大挑战。目前可解释AI主要针对自然语言处理和计算机视觉,应用于机器人时存在两大不足:缺乏对决策任务的语义锚定,且无法评估解释的可信度。本文提出一种基于人类可理解的高层概念的可信可解释机器人技术,通过将神经网络激活与人类可理解的可视化内容匹配,生成带有不确定性评分的解释。我们在多种模拟和真实机器人决策模型上进行了实验,验证了该方法作为后处理、人性化机器人诊断工具的有效性。
原文摘要 · Abstract (English)
Black box neural networks are an indispensable part of modern robots. Nevertheless, deploying such high-stakes systems in real-world scenarios poses significant challenges when the stakeholders, such as engineers and legislative bodies, lack insights into the neural networks' decision-making process. Presently, explainable AI is primarily tailored to natural language processing and computer vision, falling short in two critical aspects when applied in robots: grounding in decision-making tasks and the ability to assess trustworthiness of their explanations. In this paper, we introduce a trustworthy explainable robotics technique based on human-interpretable, high-level concepts that attribute to the decisions made by the neural network. Our proposed technique provides explanations with associated uncertainty scores for the explanation by matching neural network's activations with human-interpretable visualizations. To validate our approach, we conducted a series of experiments with various simulated and real-world robot decision-making models, demonstrating the effectiveness of the proposed approach as a post-hoc, human-friendly robot diagnostic tool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。