机器人在动态环境中如何智能安排重规划时机,以节省计算资源。
When Should a Robot Replan? Regret-Guided Update Scheduling in Time-Varying MDPs

- 采用跳步更新机制,按需重估环境动态并生成新策略。
- 理论证明更新间隔越长,后悔值增长越快,可指导预算分配。
- 适用于资源受限的移动机器人,如火星车与无人机导航任务。
在非平稳环境中运行的机器人必须持续适应环境动态变化,但机载能源与计算预算限制了全状态估计和重规划的频率。这引出一个问题:在时间轴上何时使用有限的计算资源?本文在已知转移漂移速率上界的时变马尔可夫决策过程(TVMDP)中形式化该问题。将执行建模为一种跳步更新机制:在选定更新时刻,代理通过最大似然估计转移核并计算有限时域策略;在更新之间则基于传播的状态估计复用当前策略。分析该机制的动态后悔值,揭示其在跳步间隔内的增长规律,取决于TVMDP特性与跳步长度。由此导出一种在线、基于后悔值的更新规则,实现预算自适应分配。在模拟火星车导航任务(含时变滑移动态)和室内障碍物场景下的Crazyflie四旋翼飞行器上评估,自适应分配策略优于其他预算约束基线。
原文摘要 · Abstract (English)
Robots operating in non-stationary environments must continually adapt their policies as the dynamics drift, but onboard energy and compute budgets cap how often a full state estimation and re-planning step can be performed. This raises a question: \emph{when}, along a horizon, should a robot spend its limited budget? We formulate this problem in time-varying Markov decision processes (TVMDPs) with a known bound on the rate of transition drift. We model execution as a \emph{skip-update} scheme in which, at chosen update times, the agent estimates the transition kernel by maximum likelihood and computes a finite-horizon policy, and between updates reuses this policy under a propagated state estimate. We analyze the dynamic regret of this scheme and show how it grows during skip intervals in terms of the properties of the TVMDP and the skip lengths; the resulting bound answers the opening question via an online, regret-guided update rule that allocates the budget adaptively. We evaluate the rule in a simulated Mars-rover navigation task with time-varying slip dynamics and on a Crazyflie quadrotor in indoor obstacle fields. Adaptive allocation outperforms other budgeted baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。