用逻辑约束引导粒子优化,高效生成满足复杂时序要求的机器人轨迹。
STL-SVPIO: Signal Temporal Logic guided Stein Variational Path Integral Optimization
- 将时序逻辑转为可微奖励,通过粒子迁移寻找高鲁棒性控制路径。
- 在长时序任务中成功率超基线方法,多智能体同步与排队任务也能解决。
- 适用于非线性动力学系统,如7自由度操作和半猎豹后空翻。
信号时序逻辑(STL)能够形式化表达机器人任务规划中的复杂时空约束。然而,从复杂的STL规范中合成长时间跨度的连续控制轨迹具有根本性挑战,源于STL鲁棒性目标的嵌套结构。现有基于求解器的方法(如混合整数线性规划,MILP)存在指数级扩展问题,而采样方法(如模型预测路径积分控制,MPPI)则难以处理稀疏、长时程代价。本文提出信号时序逻辑引导的斯坦因变分路径积分优化(STL-SVPIO),将STL重构为全局信息丰富且可微的奖励塑造机制。借助斯坦因变分梯度下降与可微物理引擎,STL-SVPIO使一组相互排斥的控制粒子向高鲁棒性区域迁移。该方法将稀疏逻辑满足转化为可处理的变分推断,有效规避了标准梯度方法的严重局部极小陷阱。实验表明,STL-SVPIO在传统STL任务中显著优于现有方法,兼具更高鲁棒性与效率;此外,其能解决复杂长时程任务,包括需同步与排队的多智能体协作,而基线方法或无法发现可行解,或计算不可行。最后,我们在含非线性动力学的敏捷机器人运动规划任务中验证算法泛化能力,如7-DoF机械臂操作与半猎豹后空翻。
原文摘要 · Abstract (English)
Signal Temporal Logic (STL) enables formal specification of complex spatiotemporal constraints for robotic task planning. However, synthesizing long-horizon continuous control trajectories from complex STL specifications is fundamentally challenging due to the nested structure of STL robustness objectives. Existing solver-based methods, such as Mixed-Integer Linear Programming (MILP), suffer from exponential scaling, whereas sampling methods, such as Model-Predictive Path Integral control (MPPI), struggle with sparse, long-horizon costs. We introduce Signal Temporal Logic guided Stein Variational Path Integral Optimization (STL-SVPIO), which reframes STL as a globally informative, differentiable reward-shaping mechanism. By leveraging Stein Variational Gradient Descent and differentiable physics engines, STL-SVPIO transports a mutually repulsive swarm of control particles toward high robustness regions. Our method transforms sparse logical satisfaction into tractable variational inference, mitigating the severe local minima traps of standard gradient-based methods. We demonstrate that STL-SVPIO significantly outperforms existing methods in both robustness and efficiency for traditional STL tasks. Moreover, it solves complex long-horizon tasks, including multi-agent coordination with synchronization and queuing while baselines either fail to discover feasible solutions, or become computationally intractable. Finally, we use STL-SVPIO in agile robotic motion planning tasks with nonlinear dynamics, such as 7-DoF manipulation and half cheetah back flips to show the generalizability of our algorithm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。