arXiv:2604.08983cs.RO2026-04被引 2

让机器人看懂装配手册与3D结构,精准预测零件空间姿态

AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly

论文配图:AssemLM: A Spatial Reasoning Multimodal Large Language Model for Robotic Assembly
图 1 · 摘自论文原文
  • 融合手册、点云和文本,用专用编码器提取三维几何特征
  • 在超90万样本的AssemBench上实现6维姿态推理领先性能
  • 适合做精密装配的机器人系统研发者参考

空间推理是具身智能的基础能力,尤其在机器人装配等精细操作任务中至关重要。现有基于视觉语言模型的方法主要依赖粗粒度2D感知,难以处理复杂的3D几何关系。为此,我们提出AssemLM——一种面向机器人装配的空间多模态大模型,通过整合装配手册、点云数据和文本指令,实现对关键6维装配姿态的精准预测,并具备显式的几何理解能力。为连接原始3D感知与高层语义推理,AssemLM采用专用点云编码器,提取细粒度的几何与旋转特征,支持高精度3D空间推理。同时,我们构建了AssemBench——一个包含超过90万组多模态样本、带有精确6维姿态标注的大规模装配导向空间推理基准,将评估从2D定位拓展至完整3D几何推断。大量实验与真实机器人测试表明,AssemLM在6维姿态推理上达到当前最优表现,可有效支持真实场景中的细粒度、多步骤装配任务。代码、模型及AssemBench数据集将公开发布。

原文摘要 · Abstract (English)

Spatial reasoning is a fundamental capability for embodied intelligence, especially for fine-grained manipulation tasks such as robotic assembly. Recent methods based on vision-language models (VLMs) largely rely on coarse 2D perception and struggle to perform accurate reasoning over complex 3D geometry. To address this limitation, we propose AssemLM, a spatial multimodal large language model for robotic assembly that integrates assembly manuals, point clouds, and textual instructions to predict task-critical 6D assembly poses with explicit geometric understanding. To bridge raw 3D perception and high-level linguistic reasoning, AssemLM employs a specialized point cloud encoder to extract fine-grained geometric and rotational features for accurate 3D spatial reasoning in assembly tasks. In addition, we introduce AssemBench, a large-scale benchmark for assembly-oriented spatial reasoning with over 900K multimodal samples and precise 6D pose annotations, extending evaluation from 2D grounding to full 3D geometric inference. Extensive experiments and real-robot evaluations demonstrate that AssemLM achieves state-of-the-art 6D pose reasoning performance and effectively supports fine-grained, multi-step assembly tasks in real-world settings. Code, models, and the AssemBench dataset will be made publicly available.

机器人装配空间推理多模态大模型6D姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。