arXiv:2506.12366cs.AI2025-06被引 1

用增强现实可视化智能体失败轨迹,让错误变学习资源

Ghost Policies: A New Paradigm for Understanding and Learning from Failure in Deep Reinforcement Learning

  • 通过增强现实将历史失败策略以半透明幽灵形式呈现
  • 提出人类可干预的失败研究协议,实现系统性故障分析
  • 构建人机共学闭环,推动失败可视化学习新范式

深度强化学习(DRL)智能体常表现出复杂且难以理解的失败模式,导致调试困难、难以复现,阻碍其在真实场景中的可靠部署。为解决这一关键问题,本文提出「幽灵策略」概念,并通过名为 Arvolution 的新型增强现实(AR)框架实现。Arvolution 将智能体的历史失败策略轨迹以半透明「幽灵」形式实时渲染,与当前活跃智能体在时空上共存,实现策略分歧的直观可视化。该框架独特融合:(1) 幽灵策略的 AR 可视化;(2) DRL 失调行为的分类体系;(3) 可系统性人为干扰的失败研究协议;(4) 人与智能体双向学习的双重循环机制。本文提出范式转变:将原本晦涩昂贵的失败转化为可读、可学、可利用的宝贵学习资源,为「失败可视化学习」这一新研究领域奠定基础。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (DRL) agents often exhibit intricate failure modes that are difficult to understand, debug, and learn from. This opacity hinders their reliable deployment in real-world applications. To address this critical gap, we introduce ``Ghost Policies,'' a concept materialized through Arvolution, a novel Augmented Reality (AR) framework. Arvolution renders an agent's historical failed policy trajectories as semi-transparent ``ghosts'' that coexist spatially and temporally with the active agent, enabling an intuitive visualization of policy divergence. Arvolution uniquely integrates: (1) AR visualization of ghost policies, (2) a behavioural taxonomy of DRL maladaptation, (3) a protocol for systematic human disruption to scientifically study failure, and (4) a dual-learning loop where both humans and agents learn from these visualized failures. We propose a paradigm shift, transforming DRL agent failures from opaque, costly errors into invaluable, actionable learning resources, laying the groundwork for a new research field: ``Failure Visualization Learning.''

强化学习失败分析增强现实人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。