arXiv:2503.08241cs.AIcs.CV2025-03ICLR被引 8

首个聚焦视觉导航安全的强化学习基准,测试智能体在复杂场景中的决策能力。

HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents

  • 构建三难度层级、双动作空间的视觉感知任务集
  • 实测表明现有方法存在显著奖励与成本权衡问题
  • 支持热力图可视化训练过程,适合研究安全强化学习者

推进安全自主系统需可靠基准来评估性能、分析方法并衡量智能体能力。人类主要依赖具身视觉感知以安全地导航和交互,这为强化学习(RL)智能体提供了重要参考。然而,现有基于视觉的3D基准仅涵盖简单导航任务。为此,我们提出 extbf{HASARD},一个包含多样化且复杂任务的基准,要求智能体在 extbf{D}oom 环境中进行战略决策、理解空间关系并预测短期未来。HASARD 设有三个难度等级和两种动作空间。对主流基线方法的实证评估揭示了该基准的复杂性、独特挑战及奖励-成本权衡。通过自顶向下热力图可视化训练中的导航路径,可洞察方法的学习过程。分阶段按难度递增训练提供隐式学习课程。HASARD 是首个专攻以我为中心视觉学习的安全强化学习基准,为探索当前与未来安全强化学习方法的潜力与边界提供了低成本、高洞察力的途径。环境与基线实现已开源:https://sites.google.com/view/hasard-bench/

原文摘要 · Abstract (English)

Advancing safe autonomous systems through reinforcement learning (RL) requires robust benchmarks to evaluate performance, analyze methods, and assess agent competencies. Humans primarily rely on embodied visual perception to safely navigate and interact with their surroundings, making it a valuable capability for RL agents. However, existing vision-based 3D benchmarks only consider simple navigation tasks. To address this shortcoming, we introduce \textbf{HASARD}, a suite of diverse and complex tasks to $\textbf{HA}$rness $\textbf{SA}$fe $\textbf{R}$L with $\textbf{D}$oom, requiring strategic decision-making, comprehending spatial relationships, and predicting the short-term future. HASARD features three difficulty levels and two action spaces. An empirical evaluation of popular baseline methods demonstrates the benchmark's complexity, unique challenges, and reward-cost trade-offs. Visualizing agent navigation during training with top-down heatmaps provides insight into a method's learning process. Incrementally training across difficulty levels offers an implicit learning curriculum. HASARD is the first safe RL benchmark to exclusively target egocentric vision-based learning, offering a cost-effective and insightful way to explore the potential and boundaries of current and future safe RL methods. The environments and baseline implementations are open-sourced at https://sites.google.com/view/hasard-bench/.

强化学习视觉导航安全控制基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。