arXiv:2602.08373cs.AIcs.LG2026-02中稿 · ICLR被引 2

让AI规划器学会自我纠错,实现可验证的安全决策。

Grounding Generative Planners in Verifiable Logic: A Hybrid Architecture for Trustworthy Embodied AI

  • 用逻辑导师与大模型对话,主动修正错误计划。
  • 在家庭安全任务中零危险动作率,77.3%目标达成率。
  • 适合追求高可靠性的机器人与自动驾驶系统研究者。

大型语言模型(LLMs)在具身智能规划中展现出潜力,但其随机性缺乏形式化推理,无法为物理部署提供严格安全保证。现有方法通常依赖不可靠的LLM进行安全检查,或仅拒绝不安全计划而不提供修复方案。我们提出可验证迭代优化框架(VIRF),一种神经符号架构,将安全机制从被动守门转向主动协作。核心创新在于一个基于形式化安全本体的确定性逻辑导师,为LLM规划器提供因果与教学式反馈,实现智能计划修复而非简单规避。我们还设计了可扩展的知识获取流程,从真实文档中合成安全知识库,弥补现有基准的盲点。在复杂家庭安全任务中,VIRF实现0%危险动作率(HAR)和77.3%目标条件率(GCR),优于所有基线,平均仅需1.1次修正迭代。VIRF为构建根本可信、可验证安全的具身智能体提供了系统性路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show promise as planners for embodied AI, but their stochastic nature lacks formal reasoning, preventing strict safety guarantees for physical deployment. Current approaches often rely on unreliable LLMs for safety checks or simply reject unsafe plans without offering repairs. We introduce the Verifiable Iterative Refinement Framework (VIRF), a neuro-symbolic architecture that shifts the paradigm from passive safety gatekeeping to active collaboration. Our core contribution is a tutor-apprentice dialogue where a deterministic Logic Tutor, grounded in a formal safety ontology, provides causal and pedagogical feedback to an LLM planner. This enables intelligent plan repairs rather than mere avoidance. We also introduce a scalable knowledge acquisition pipeline that synthesizes safety knowledge bases from real-world documents, correcting blind spots in existing benchmarks. In challenging home safety tasks, VIRF achieves a perfect 0 percent Hazardous Action Rate (HAR) and a 77.3 percent Goal-Condition Rate (GCR), which is the highest among all baselines. It is highly efficient, requiring only 1.1 correction iterations on average. VIRF demonstrates a principled pathway toward building fundamentally trustworthy and verifiably safe embodied agents.

具身智能安全规划逻辑推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。