arXiv:2501.01463cs.LGcs.AI2025-01被引 5

用深度强化学习从原始数据中自动识别目标,效果更好且更省资源。

Goal Recognition using Actor-Critic Optimization

  • 基于演员-评论家框架,直接从无结构数据学策略网络进行推理
  • 在离散与连续场景下均达当前最佳性能,计算和内存成本显著降低
  • 适合需要自动建模复杂行为的目标识别任务

目标识别旨在从观测序列中推断智能体的目标。现有方法通常依赖人工设计的领域和离散表示。深度目标识别(DRACO)是一种基于深度强化学习的新方法,克服了这些限制,主要贡献有两点:一是首个从非结构化数据中学习策略网络并用于推理的目标识别算法;二是引入新度量方式,通过连续策略表示评估目标假设。DRACO在离散设置下达到当前最优性能,且不依赖传统结构化输入;在更具挑战性的连续设置中表现优于现有方法,同时大幅降低计算与内存开销。这些结果展示了新算法的鲁棒性,成功连接了传统目标识别与深度强化学习。

原文摘要 · Abstract (English)

Goal Recognition aims to infer an agent's goal from a sequence of observations. Existing approaches often rely on manually engineered domains and discrete representations. Deep Recognition using Actor-Critic Optimization (DRACO) is a novel approach based on deep reinforcement learning that overcomes these limitations by providing two key contributions. First, it is the first goal recognition algorithm that learns a set of policy networks from unstructured data and uses them for inference. Second, DRACO introduces new metrics for assessing goal hypotheses through continuous policy representations. DRACO achieves state-of-the-art performance for goal recognition in discrete settings while not using the structured inputs used by existing approaches. Moreover, it outperforms these approaches in more challenging, continuous settings at substantially reduced costs in both computing and memory. Together, these results showcase the robustness of the new algorithm, bridging traditional goal recognition and deep reinforcement learning.

目标识别强化学习策略网络连续表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。