arXiv:2509.24163cs.RO2025-09

用多模态大模型推理物体堆叠偏好,提升长时序机械臂操作成功率。

Preference-Based Long-Horizon Robotic Stacking with Multimodal Large Language Models

  • 基于多模态输入和堆叠偏好推理生成最优堆叠顺序。
  • 在仿真中相比基线模型堆叠完成率显著提升。
  • 适用于真实人形机器人在线执行复杂堆叠任务。

预训练的大语言模型(LLM)可通过推理抽象任务描述和自然语言指令作为高层机器人规划器,但在涉及物体物理属性的长时序机械操作任务中表现不足。例如,内部藏有物品的容器堆叠需考虑重量、稳定性等隐藏物理特性。为此,本文提出使用多模态大语言模型作为高层规划器,对每个待堆叠物体进行多模态输入,并通过推理堆叠偏好确定最佳序列。为使模型能同时处理多种偏好而不依赖显式指令,我们构建了一个包含重量、稳定性、尺寸和占地面积等偏好的定制数据集以微调模型。大规模仿真评估表明,相比提示调优的预训练模型,该方法显著提升了堆叠完成率。此外,我们在真实人形机器人上实现了该框架的在线应用,验证了其有效性。

原文摘要 · Abstract (English)

Pretrained large language models (LLMs) can work as high-level robotic planners by reasoning over abstract task descriptions and natural language instructions, etc. However, they have shown a lack of knowledge and effectiveness in planning long-horizon robotic manipulation tasks where the physical properties of the objects are essential. An example is the stacking of containers with hidden objects inside, which involves reasoning over hidden physics properties such as weight and stability. To this end, this paper proposes to use multimodal LLMs as high-level planners for such long-horizon robotic stacking tasks. The LLM takes multimodal inputs for each object to stack and infers the current best stacking sequence by reasoning over stacking preferences. Furthermore, in order to enable the LLM to reason over multiple preferences at the same time without giving explicit instructions, we propose to create a custom dataset considering stacking preferences including weight, stability, size, and footprint, to fine-tune the LLM. Compared to the pretrained LLM with prompt tuning, we demonstrate the improved stacking completion of the LLM fine-tuned with our custom dataset via large-scale simulation evaluation. Furthermore, we showcase the effectiveness of the proposed framework for the long-horizon stacking task on a real humanoid robot in an online manner.

机器人规划多模态大模型堆叠任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。