解决服务机器人在动态不确定奖励下的路径规划难题
Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics

- 基于观测预测未来奖励,实现在线自适应路径规划
- 长时域规划结合实时调整,显著提升任务完成率
- 专为室内服务机器人设计的基准测试平台
我们提出不确定时变奖励下的定向问题(OP-UTVR),这是经典定向问题(OP)的新变体。传统OP假设奖励已知,但实际应用中如配送员面对波动的客户需求,奖励具有不确定性且随时间变化。OP-UTVR允许代理通过观测估计奖励动态并预测未来收益,从而在随机变化和预测误差下做出合理路径决策。本文设计了三种不同规划周期与在线适应能力的规划器,并推导了其在奖励随机性下的性能理论边界。此外,我们构建了一个面向移动服务机器人的OP-UTVR基准测试,机器人需在室内人流环境中导航。实验揭示了规划周期与适应性间的权衡,验证了结合长期规划与在线调整的有效性。
原文摘要 · Abstract (English)
We present the orienteering problem with uncertain time-varying rewards (OP-UTVR), a novel variant of the orienteering problem (OP). While most existing OP formulations assume rewards to be known in advance, practical applications involve uncertain and time-varying rewards, as with shifting customer demand for delivery agents. OP-UTVR relaxes this assumption by allowing agents to estimate reward dynamics from observations and forecast future rewards. This enables informed routing decisions despite stochastic reward changes and inevitable prediction errors. We address this problem using three planners that differ in planning horizon and online adaptivity, and derive theoretical bounds on their performance under reward stochasticity. We further introduce a mobile service robot benchmark for OP-UTVR, where a robot navigates among pedestrians in indoor environments. Experiments reveal trade-offs between planning horizon and adaptivity, and demonstrate the effectiveness of long-horizon planning with online adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。