让机器人读懂人类赛车意图,实时辅助实现协同竞速
Dreaming to Assist: Learning to Align with Human Objectives for Shared Control in High-Speed Racing
- 用递归状态空间模型推断人类战术目标
- 混合人机动作后胜过纯虚拟对手和基线策略
- 适合高速动态协作场景,如赛车、无人机编队
在高速动态且需战术决策的多车竞速场景中,人机团队需要紧密协作。为解决这一问题,我们提出 Dream2Assist 框架,结合一个能推断人类目标与价值函数的丰富世界模型,以及一个可提供适配专家协助的辅助智能体。该方法基于递归状态空间模型显式推断人类意图,使辅助智能体能够选择与人类目标一致的动作,实现流畅的人机协同。我们在包含多种合成人类司机(如“保持后方”和“超车”)的高速竞速环境中验证了该方法。结果表明,融合人机动作的团队表现优于纯合成人类对手及多个基线辅助策略,且意图条件控制确保任务执行中遵循人类偏好,提升性能同时满足其目标。
原文摘要 · Abstract (English)
Tight coordination is required for effective human-robot teams in domains involving fast dynamics and tactical decisions, such as multi-car racing. In such settings, robot teammates must react to cues of a human teammate's tactical objective to assist in a way that is consistent with the objective (e.g., navigating left or right around an obstacle). To address this challenge, we present Dream2Assist, a framework that combines a rich world model able to infer human objectives and value functions, and an assistive agent that provides appropriate expert assistance to a given human teammate. Our approach builds on a recurrent state space model to explicitly infer human intents, enabling the assistive agent to select actions that align with the human and enabling a fluid teaming interaction. We demonstrate our approach in a high-speed racing domain with a population of synthetic human drivers pursuing mutually exclusive objectives, such as "stay-behind" and "overtake". We show that the combined human-robot team, when blending its actions with those of the human, outperforms the synthetic humans alone as well as several baseline assistance strategies, and that intent-conditioning enables adherence to human preferences during task execution, leading to improved performance while satisfying the human's objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。