arXiv:2602.21161cs.RO2026-02中稿 · the 2026 IEEE Inte…被引 1

用大模型推理物理动作,让机器人更聪明地堆叠积木

ActionReasoning: Robot Action Reasoning in 3D Space with LLM for Robotic Brick Stacking

  • 通过多智能体大模型进行物理感知的动作规划
  • 实现稳定积木堆叠,减少底层代码依赖
  • 适合希望提升机器人通用性的研究者和开发者

传统机器人系统依赖为特定环境设计的定制规划器,虽在受限场景有效,但泛化能力差,限制了具身智能与通用机器人的发展。近期基于数据的视觉-语言-动作(VLA)方法试图从大规模仿真和真实世界数据中学习策略,但物理世界的连续动作空间远超语言标记的表征能力,仅靠数据扩展难以实现通用机器人智能。为此,我们提出ActionReasoning:一种基于大语言模型(LLM)的框架,通过显式动作推理生成符合物理规律、基于先验知识的决策。该框架利用已编码于LLM中的物理先验与现实知识,构建多智能体架构。我们在积木堆叠这一具体案例中验证该方法,假设环境状态可准确测量,将状态序列化后输入多智能体LLM框架,生成具备物理意识的动作计划。实验表明,该框架能实现稳定的积木放置,将工作重心从低层领域特定编码转向高层工具调用与提示工程,展现出良好的泛化潜力。本工作为融合感知与执行提供了新路径,推动机器人操作中物理推理与大模型的结合。

原文摘要 · Abstract (English)

Classical robotic systems typically rely on custom planners designed for constrained environments. While effective in restricted settings, these systems lack generalization capabilities, limiting the scalability of embodied AI and general-purpose robots. Recent data-driven Vision-Language-Action (VLA) approaches aim to learn policies from large-scale simulation and real-world data. However, the continuous action space of the physical world significantly exceeds the representational capacity of linguistic tokens, making it unclear if scaling data alone can yield general robotic intelligence. To address this gap, we propose ActionReasoning, an LLM-driven framework that performs explicit action reasoning to produce physics-consistent, prior-guided decisions for robotic manipulation. ActionReasoning leverages the physical priors and real-world knowledge already encoded in Large Language Models (LLMs) and structures them within a multi-agent architecture. We instantiate this framework on a tractable case study of brick stacking, where the environment states are assumed to be already accurately measured. The environmental states are then serialized and passed to a multi-agent LLM framework that generates physics-aware action plans. The experiments demonstrate that the proposed multi-agent LLM framework enables stable brick placement while shifting effort from low-level domain-specific coding to high-level tool invocation and prompting, highlighting its potential for broader generalization. This work introduces a promising approach to bridging perception and execution in robotic manipulation by integrating physical reasoning with LLMs.

机器人大模型物理推理积木堆叠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。