arXiv:2603.11461cs.RO2026-03中稿 · ASME MSEC 2026

用大模型让机器人和人协作组装新零件,不依赖预设流程。

CoViLLM: An Adaptive Human-Robot Collaborative Assembly Framework Using Large Language Models

  • 结合深度相机定位、人体操作识别和大模型规划任务
  • 在3类产品上实现灵活装配,无需预设流程
  • 适合需要快速切换产品的智能制造场景

随着大规模定制需求增加,传统基于规则的制造机器人难以适应个性化或新产品变体。人机协作可通过发挥人类灵活性与决策能力提升系统适应性。然而,现有框架通常依赖预设的感知-操作流水线,无法自主生成新产品的任务计划。本文提出CoViLLM,一种支持定制化及未见过产品装配的自适应人机协同装配框架。该框架融合基于深度相机的定位以估计物体位置,通过人类操作员分类识别新组件,并利用大语言模型根据自然语言指令进行装配任务规划。在NIST装配任务板上对已知、定制化及新产品案例进行了验证。实验结果表明,所提框架能通过扩展人机协作的边界,在非预设产品与任务设置下实现灵活装配。

原文摘要 · Abstract (English)

With increasing demand for mass customization, traditional manufacturing robots that rely on rule-based operations lack the flexibility to accommodate customized or new product variants. Human-Robot Collaboration has demonstrated potential to improve system adaptability by leveraging human versatility and decision-making capabilities. However, existing Human-Robot Collaborative frameworks typically depend on predefined perception-manipulation pipelines, limiting their ability to autonomously generate task plans for new product assembly. In this work, we propose CoViLLM, an adaptive human-robot collaborative assembly framework that supports the assembly of customized and previously unseen products. CoViLLM combines depth-camera-based localization for object position estimation, human operator classification for identifying new components, and a Large Language Model for assembly task planning based on natural language instructions. The framework is validated on the NIST Assembly Task Board for known, customized, and new product cases. Experimental results show that the proposed framework enables flexible collaborative assembly by extending Human-Robot Collaboration beyond predefined product and task settings.

人机协作大模型智能制造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。