让机器人学会根据时间调整行为,提升灵活性与适应性。
Time as a Control Dimension in Robot Learning
- 将时间作为可调控变量,动态调整策略执行节奏。
- 在多种任务中实现从快速执行到精细操作的连续过渡。
- 无需重训练即可应对时间变化、干扰和人工干预。
时间感知在智能行为中起核心作用,影响动作的节奏、协调与对变化目标和环境的适应。然而,多数机器人学习算法仅将时间视为固定的回合周期或调度约束。本文提出时间感知策略学习框架,将时间作为机器人行为的可控维度。该方法为策略引入两个时间信号:剩余时间和时间比例,后者调节策略内部的时间推进,使单一策略能在不同时间尺度下灵活调整执行策略。在长时程操作、颗粒介质倾倒、柔性物体交互及多智能体协同等多样化任务中,策略能持续适应从紧张时限下的动态执行到有充裕时间时的稳定、谨慎交互。该方法提升了效率、在仿真到现实迁移中的鲁棒性以及面对扰动和人工输入时的可控性,且无需重新训练。将时间视为可控变量,为自适应与人机对齐的机器人自主提供了新范式。
原文摘要 · Abstract (English)
Temporal awareness plays a central role in intelligent behavior by shaping how actions are paced, coordinated, and adapted to changing goals and environments. In contrast, most robot learning algorithms treat time only as a fixed episode horizon or scheduling constraint. Here we introduce time-aware policy learning, a reinforcement learning framework that treats time as a control dimension of robot behavior. The approach augments policies with two temporal signals, the remaining time and a time ratio that modulates the policy's internal progression of time, allowing a single policy to regulate its execution strategy across temporal regimes. Across diverse manipulation tasks including long-horizon manipulation, granular-media pouring, articulated-object interaction, and multi-agent coordination, the resulting policies adapt their behavior continuously from dynamic execution under tight schedules to stable and deliberate interaction when more time is available. This temporal awareness improves efficiency, robustness under sim-to-real mismatch and disturbances, and controllability through human input without retraining. Treating time as a controllable variable provides a new framework for adaptive and human-aligned robot autonomy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。