arXiv:2512.17321cs.RO2025-12被引 3

用符号推理+神经控制结合,让语言指令操控机器人更稳更快。

Neuro-Symbolic Control with Large Language Models for Language-Guided Spatial Tasks

  • 分层设计:语言模型处理语义,神经控制器执行动作
  • 成功率提升,平均步数减少70%以上,速度最高快8.83倍
  • 无需强化学习,适合追求稳定性的实际应用

尽管大语言模型(LLMs)在具身系统中已展现出语言条件控制的潜力,但其不稳定性、收敛慢及幻觉动作仍限制了在连续控制中的直接应用。本文提出一种模块化的神经符号控制框架,明确区分低层运动执行与高层语义推理。轻量级神经增量控制器在连续空间中执行有界、逐步的动作,而本地部署的LLM负责解释符号化任务。在平面操作场景中,通过语言描述物体间的空间关系进行评估。大量实验使用Mistral、Phi和LLaMA-3.2等本地语言模型,对比了纯LLM控制、纯神经控制与提出的LLM+DL框架。结果表明,相较于纯LLM基线,神经符号融合显著提升了成功率与效率,平均步数减少超70%,速度最高提升8.83倍,且对语言模型质量不敏感。该框架通过将LLM输出限定为符号化指令,并将未解释的执行交由在人工几何数据上训练的神经控制器完成,实现了可解释性、稳定性与泛化性的提升,无需强化学习或昂贵的试错。实验证明,神经符号分解为语言理解与持续控制的集成提供了一种可扩展且原理清晰的路径,推动可靠高效的语言引导具身系统发展。

原文摘要 · Abstract (English)

Although large language models (LLMs) have recently become effective tools for language-conditioned control in embodied systems, instability, slow convergence, and hallucinated actions continue to limit their direct application to continuous control. A modular neuro-symbolic control framework that clearly distinguishes between low-level motion execution and high-level semantic reasoning is proposed in this work. While a lightweight neural delta controller performs bounded, incremental actions in continuous space, a locally deployed LLM interprets symbolic tasks. We assess the suggested method in a planar manipulation setting with spatial relations between objects specified by language. Numerous tasks and local language models, such as Mistral, Phi, and LLaMA-3.2, are used in extensive experiments to compare LLM-only control, neural-only control, and the suggested LLM+DL framework. In comparison to LLM-only baselines, the results show that the neuro-symbolic integration consistently increases both success rate and efficiency, achieving average step reductions exceeding 70% and speedups of up to 8.83x while remaining robust to language model quality. The suggested framework enhances interpretability, stability, and generalization without any need of reinforcement learning or costly rollouts by controlling the LLM to symbolic outputs and allocating uninterpreted execution to a neural controller trained on artificial geometric data. These outputs show empirically that neuro-symbolic decomposition offers a scalable and principled way to integrate language understanding with ongoing control, this approach promotes the creation of dependable and effective language-guided embodied systems.

语言控制神经符号机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。