arXiv:2609.05190cs.AI2026-09被引 2

提出可解释智能体行为的镜像模型架构

The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior

论文配图:The Mirror Agent Model: a Bayesian Architecture for Interpretable Agent Behavior
图 1 · 摘自论文原文
  • 用镜像机制构建可解释的行为生成框架
  • 通过现成显著性方法实现行为解释可视化
  • 适合需要透明决策过程的研究与应用

本文提出一种新型架构,用于生成可解释的行为与解释。该架构被称为镜像智能体模型(Mirror Agent Model),其核心思想是将观察者模型设计为智能体自身的镜像,以实现显式与隐式沟通的目标。为提供整体理解,我们首先回顾了相关工作中关于智能体意图传达与可读行为生成的研究进展。随后,论文引入基于现成显著性方法的新解释能力,并展示初步的定性结果,验证了该架构在生成可解释行为方面的潜力。

原文摘要 · Abstract (English)

In this paper we illustrate a novel architecture generating interpretable behavior and explanations. We refer to this architecture as the Mirror Agent Model because it defines the observer model, that is the target of explicit and implicit communications, as a mirror of the agent's. With the goal of providing a general understanding of this work, we firstly show prior relevant results addressing the informative communication of agents intentions and the production of legible behavior. In the second part of the paper we furnish the architecture with novel capabilities for explanations through off-the-shelf saliency methods, followed by preliminary qualitative results.

可解释智能体贝叶斯建模行为解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。