用形式化逻辑提升视觉语言模型的指令执行精度
STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models

- 用信号时序逻辑统一语言理解与机器人执行
- 实测在桌面上提升指令执行的精准度与可靠性
- 适合需要高可靠性的智能机器人任务设计
视觉-语言-动作(VLA)模型虽具强大泛化能力,但常缺乏可解释性,难以遵循包含空间、时间及逻辑要求的精确自然语言指令。本文提出一种分层框架,采用信号时序逻辑(STL)作为连接高层语言理解与底层机器人执行的共享表示。高层策略利用视觉语言模型将语言指令分解为子任务,生成对应STL规范,并选择低层策略执行每个子任务。STL规范将语言意图转化为精确约束,而低层策略选择决定这些约束是否通过STL引导的模型预测控制直接执行,或在学习策略执行中进行监控,以应对感知复杂或接触密集的行为。通过在计划验证、低层策略、子任务监控和重规划中集成STL,本框架实现了基于语言的计划在运行时的检查、优化与修正。我们在真实桌面场景中评估该方法,证明形式化规范能显著提升语言驱动机器人规划的精度、可靠性和可解释性。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have shown impressive generalization, but often lack interpretability and can struggle to follow precise natural language instructions that encode spatial, temporal, and logical requirements. We propose a hierarchical framework that uses Signal Temporal Logic (STL) as a shared representation connecting high-level language understanding with low-level robot execution. A high-level policy leverages a VLM to decompose language instructions into high-level subtasks, generate STL specifications for each subtask, and choose a low-level policy for executing each subtask. The STL specifications translate language-derived intent into precise constraints, and the low-level policy selection determines whether those constraints are enforced directly through STL-guided model-predictive control or monitored during execution of a learned policy for perceptually complex, or contact-rich behaviors. By integrating STL into plan validation, low-level policy, subtask monitoring, and replanning, our framework enables language-derived plans to be checked, optimized, and revised at runtime using a common formal structure. We evaluate the approach on a real-world tabletop domain, demonstrating how formal specifications can improve the precision, reliability, and interpretability of language-conditioned robot planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。