arXiv:2501.18086cs.LGcs.AI2025-01

用共享约束分布自适应学习,让机器人在多任务中自动保持安全。

DIAL: Distribution-Informed Adaptive Learning of Multi-Task Constraints for Safety-Critical Systems

  • 通过模仿学习提取多任务共享的安全约束分布
  • 新任务只需调整风险水平,无需重定义约束条件
  • 适用于自动驾驶等复杂场景,适合需要高安全性的应用

安全强化学习通常依赖预设的约束函数以确保复杂现实任务(如自动驾驶)中的安全性。然而,在不同任务中准确设定这些函数仍具挑战性。近期研究指出,利用预先获取的任务无关知识可提升相关任务的安全性和样本效率。基于此,我们提出一种新方法,用于学习多个任务间的共享约束分布。该方法通过模仿学习识别共享约束,并在新任务中通过调整这些分布内的风险水平实现自适应。这种灵活性可缓解专家偏好带来的风险敏感度差异,确保即使示范不完美,仍能一致遵循通用安全原则。该方法适用于控制与导航领域,涵盖多任务及元任务场景,支持如保持安全距离或遵守限速等约束。实验验证表明,相比基线方法,本方法在无需任务特定约束定义的前提下,显著提升了安全表现与成功率,展现出在广泛现实任务中的通用性与实用性。

原文摘要 · Abstract (English)

Safe reinforcement learning has traditionally relied on predefined constraint functions to ensure safety in complex real-world tasks, such as autonomous driving. However, defining these functions accurately for varied tasks is a persistent challenge. Recent research highlights the potential of leveraging pre-acquired task-agnostic knowledge to enhance both safety and sample efficiency in related tasks. Building on this insight, we propose a novel method to learn shared constraint distributions across multiple tasks. Our approach identifies the shared constraints through imitation learning and then adapts to new tasks by adjusting risk levels within these learned distributions. This adaptability addresses variations in risk sensitivity stemming from expert-specific biases, ensuring consistent adherence to general safety principles even with imperfect demonstrations. Our method can be applied to control and navigation domains, including multi-task and meta-task scenarios, accommodating constraints such as maintaining safe distances or adhering to speed limits. Experimental results validate the efficacy of our approach, demonstrating superior safety performance and success rates compared to baselines, all without requiring task-specific constraint definitions. These findings underscore the versatility and practicality of our method across a wide range of real-world tasks.

强化学习安全控制多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。