直接从多视角视觉和逻辑任务指令生成可执行路径,解决真实环境中的规划难题。
Bridging Perception and Planning: Towards End-to-End Planning for Signal Temporal Logic Tasks
- 用可微分框架将多视角图像与STL任务直接映射为轨迹
- 在工厂物流仿真中,满足逻辑约束的轨迹成功率超基线18%
- 推理时加安全滤波器,既保逻辑正确又提升物理可执行性
我们研究机器人在信号时序逻辑(STL)规范下的任务与运动规划问题。现有STL方法依赖预定义地图或移动性表示,在非结构化真实环境中效果不佳。本文提出结构化混合专家型STL规划器(S-MSP),一种可微分框架,能将同步多视角相机观测与STL规范直接映射为可行轨迹。S-MSP在统一流程中集成STL约束,采用包含轨迹重建与STL鲁棒性的复合损失进行训练。采用结构感知的混合专家(MoE)模型,通过时间锚定嵌入实现时域任务分解与专用化。我们在高保真工厂物流场景下评估S-MSP,实验表明其在满足STL规范与轨迹可行性上优于单专家基线。推理阶段引入基于规则的安全滤波器,显著提升物理可执行性,且不损害逻辑正确性,验证了方法的实用性。
原文摘要 · Abstract (English)
We investigate the task and motion planning problem for Signal Temporal Logic (STL) specifications in robotics. Existing STL methods rely on pre-defined maps or mobility representations, which are ineffective in unstructured real-world environments. We propose the \emph{Structured-MoE STL Planner} (\textbf{S-MSP}), a differentiable framework that maps synchronized multi-view camera observations and an STL specification directly to a feasible trajectory. S-MSP integrates STL constraints within a unified pipeline, trained with a composite loss that combines trajectory reconstruction and STL robustness. A \emph{structure-aware} Mixture-of-Experts (MoE) model enables horizon-aware specialization by projecting sub-tasks into temporally anchored embeddings. We evaluate S-MSP using a high-fidelity simulation of factory-logistics scenarios with temporally constrained tasks. Experiments show that S-MSP outperforms single-expert baselines in STL satisfaction and trajectory feasibility. A rule-based \emph{safety filter} at inference improves physical executability without compromising logical correctness, showcasing the practicality of the approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。