arXiv:2510.27558cs.ROcs.AI2025-10

用现成大模型实现无需训练的长时程机器人操作。

Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs

  • 用场景图保持环境空间认知,支持持续推理。
  • 直接调用现成大模型,无需领域微调。
  • 适合想快速搭建通用机械臂系统的研究者。

本文提出一个框架,利用预训练的基础模型实现无需领域特定训练的机器人操作。该框架整合现成模型,结合基础模型的多模态感知与通用推理模型,具备鲁棒的任务序列规划能力。框架内动态维护的场景图提供空间意识,支持对环境的一致性推理。通过一系列桌面机器人操作实验评估,结果表明该框架可直接基于现成基础模型构建机器人操作系统,具有显著应用潜力。

原文摘要 · Abstract (English)

This paper presents a framework that leverages pre-trained foundation models for robotic manipulation without domain-specific training. The framework integrates off-the-shelf models, combining multimodal perception from foundation models with a general-purpose reasoning model capable of robust task sequencing. Scene graphs, dynamically maintained within the framework, provide spatial awareness and enable consistent reasoning about the environment. The framework is evaluated through a series of tabletop robotic manipulation experiments, and the results highlight its potential for building robotic manipulation systems directly on top of off-the-shelf foundation models.

机器人操作大模型场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。