arXiv:2608.20556cs.ROcs.LO2026-08被引 1

让机器人按逻辑规则执行任务,一次模型适配多种安全约束。

Logic-VLA: A Temporal Logic Conditioned Vision-Language-Action Model

论文配图:Logic-VLA: A Temporal Logic Conditioned Vision-Language-Action Model
图 1 · 摘自论文原文
  • 用时序逻辑条件控制视觉语言动作模型行为
  • 逻辑满足率提升24.8至40.7个百分点,任务成功率下降少于1.8个点
  • 适合需要高安全性和动态约束的智能体控制场景

视觉-语言-动作(VLA)模型能理解自然语言任务指令,但这类指令难以精确表达安全关键或时空要求。本文提出Logic-VLA,一种在推理时基于信号时序逻辑(STL)规范进行条件控制的正式需求感知型VLA。其采用语法图结构的STL编码器,预先训练以捕捉时序逻辑语义。策略适配分两阶段:先在满足性示范上进行STL条件监督微调,再通过流匹配代理的轨迹级偏好优化,对匹配的满足/违反回滚对进行优化。该方法在保持原始自然语言任务性能的同时显著提升形式化要求满足率。我们在随机化照片级真实感环境中评估了闭环无人机导航,并测试了对训练中未见的STL公式的泛化能力。在所有基准测试中,Logic-VLA相较无STL感知的基线策略,使STL满足率提升24.8至40.7个百分点,同时将自然语言任务成功率下降控制在最多1.8个百分点以内,证明单一VLA可在不为每种规范单独训练策略的前提下,灵活适应多样化的形式化需求。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models can follow natural-language (NL) task instructions, but such instructions may not precisely specify safety-critical or spatiotemporal requirements on the resulting behavior. We introduce Logic-VLA, a formal-requirement-aware VLA that conditions on Signal Temporal Logic (STL) specifications supplied at inference time. Logic-VLA uses a syntax-graph-based STL encoder pre-trained to capture temporal logic semantics. Policy adaptation proceeds in two stages: STL-conditioned supervised fine-tuning on satisfying demonstrations is followed by trajectory-level preference optimization over matched satisfying-violating rollout pairs using a flow-matching surrogate for Identity Preference Optimization. This formulation improves formal requirement satisfaction while preserving the nominal NL task. We evaluate Logic-VLA in closed-loop quadcopter navigation simulation across randomized photorealistic environments and test generalization to STL formulas unseen during training. Across the evaluation benchmarks, Logic-VLA improves STL satisfaction rate over an STL-blind base policy by 24.8 to 40.7 percentage points (pp) while reducing nominal NL task success by at most 1.8 pp, showing that a single VLA can adapt its behavior to varying formal requirements without requiring a separate policy for each specification.

视觉语言动作时序逻辑机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。