用大模型组合实现人机协同拆解电动车电池,指令越简单,规划越可靠。
Intent-Driven LLM Ensemble Planning for Flexible Multi-Robot Disassembly: Demonstration on EV Batteries
- 基于人类语言指令与多机器人能力匹配,生成可执行的拆解动作序列。
- 在200个真实场景中,完整序列正确率达94.7%,下一步动作正确率98.2%。
- 适合需要低人工干预的自动化回收、工业协作等场景。
本文针对多机器人在非结构化场景中执行复杂操作任务的问题,提出一种意图驱动的规划流水线。该流水线整合:(i) 视觉到文本的场景编码,(ii) 多个大语言模型(LLMs)根据操作员意图生成候选移除序列,(iii) 基于LLM的验证器确保格式与先后顺序约束,(iv) 确定性一致性过滤器剔除幻觉物体。在电动汽车电池拆解任务中评估,两个机械臂需协同完成600条不同指令下的拆解。在200个真实场景、5类组件上测试,采用完整序列正确率和下一步动作正确率作为指标,对比五种基于LLM的规划器,并进行消融分析。同时通过人类实验评估界面效率,使用执行时间与NASA TLX量表。结果表明,该组合验证方法能可靠将人类意图转化为安全可执行的多机器人计划,且用户负担低。
原文摘要 · Abstract (English)
This paper addresses the problem of planning complex manipulation tasks, in which multiple robots with different end-effectors and capabilities, informed by computer vision, must plan and execute concatenated sequences of actions on a variety of objects that can appear in arbitrary positions and configurations in unstructured scenes. We propose an intent-driven planning pipeline which can robustly construct such action sequences with varying degrees of supervisory input from a human using simple language instructions. The pipeline integrates: (i) perception-to-text scene encoding, (ii) an ensemble of large language models (LLMs) that generate candidate removal sequences based on the operator's intent, (iii) an LLM-based verifier that enforces formatting and precedence constraints, and (iv) a deterministic consistency filter that rejects hallucinated objects. The pipeline is evaluated on an example task in which two robot arms work collaboratively to dismantle an Electric Vehicle battery for recycling applications. A variety of components must be grasped and removed in specific sequences, determined by human instructions and/or by task-order feasibility decisions made by the autonomous system. On 200 real scenes with 600 operator prompts across five component classes, we used metrics of full-sequence correctness and next-task correctness to evaluate and compare five LLM-based planners (including ablation analyses of pipeline components). We also evaluated the LLM-based human interface in terms of time to execution and NASA TLX with human participant experiments. Results indicate that our ensemble-with-verification approach reliably maps operator intent to safe, executable multi-robot plans while maintaining low user effort.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。