让智能体通过动作暴露状态,便于远程监控和协作
Training Observable Control Policies to Expose Agent State Through Actions

- 用强化学习训练策略,使动作更易反映内部状态
- 在飞机追踪任务中实现高可观测性,性能损失极小
- 适合通信受限的多智能体系统设计
物理或操作约束常导致自主智能体通信受限,增加监控或多智能体协调难度。即使缺乏强通信,部分状态信息仍可通过观测智能体行为推断。本文研究利用智能体与环境交互产生的动作作为状态估计的信息源,采用强化学习训练策略,通过奖励机制增强策略的可观测性。通过仿真分析训练后的策略,在飞机追踪任务中发现一种可观测性显著提升的策略,对原任务性能影响微乎其微。
原文摘要 · Abstract (English)
Physical or operational constraints often impose communications limitations on autonomous agents. Such limitations complicate monitoring or multiagent coordination. Even when strong communications are absent, some information may still be available. The remainder of the relevant agent state may be reconstructed via estimation. The actions taken by an agent are a potential source of information -- as the agent interacts with the environment, these actions may be observed even in the absence of explicit communication. We investigate using actions to estimate the state of an agent, using reinforcement learning to develop policies which make the estimation problem more tractable. Policy observability is encouraged through the training reward and is analyzed using simulation of the trained agent. In an aircraft tracking problem a policy with enhanced observability is found that has minimal impact on nominal task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。