arXiv:2512.04446cs.ROcs.SY2025-12被引 5

用视觉语言动作模型实现电脑部件精准拆解,发现混合策略更有效。

Vision-Language-Action Models for Selective Robotic Disassembly: A Case Study on Critical Component Extraction from Desktops

  • 将拆解任务分解为小步骤,用VLA模型端到端学习
  • 微调后模型能完成早期步骤但关键操作易失败
  • 结合规则控制器的混合方法可成功完成全流程

针对报废台式机中高价值或敏感部件(如内存条、CPU、硬盘)的自动化拆解仍面临挑战,主要源于产品多样性与操作精度要求。传统拆解流程分感知、规划、运动等多个阶段,需显式建模,泛化能力差。近年来,视觉-语言-动作(VLA)模型为机器人操作提供了端到端解决方案。本文构建了专用数据集,对OpenVLA和OpenVLA-OFT进行微调,以探索其在复杂拆解任务中的适用性。实验表明,微调后的模型可准确执行多个前期步骤,但在关键子任务上表现不佳,导致整体失败。然而,采用简单混合策略——将VLA与规则控制器结合,能够顺利完成全部拆解流程。研究揭示了当前VLA模型在处理精细操作时的局限性,并为未来端到端自动化拆解提供了重要参考。

原文摘要 · Abstract (English)

Automating disassembly of critical components from end-of-life (EoL) desktops, such as high-value items like RAM modules and CPUs, as well as sensitive parts like hard disk drives, remains challenging due to the inherent variability and uncertainty of these products. Moreover, their disassembly requires sequential, precise, and dexterous operations, further increasing the complexity of automation. Current robotic disassembly processes are typically divided into several stages: perception, sequence planning, task planning, motion planning, and manipulation. Each stage requires explicit modeling, which limits generalization to unfamiliar scenarios. Recent development of vision-language-action (VLA) models has presented an end-to-end approach for general robotic manipulation tasks. Although VLAs have demonstrated promising performance on simple tasks, the feasibility of applying such models to complex disassembly remains largely unexplored. In this paper, we collected a customized dataset for robotic RAM and CPU disassembly and used it to fine-tune two well-established VLA approaches, OpenVLA and OpenVLA-OFT, as a case study. We divided the whole disassembly task into several small steps, and our preliminary experimental results indicate that the fine-tuned VLA models can faithfully complete multiple early steps but struggle with certain critical subtasks, leading to task failure. However, we observed that a simple hybrid strategy that combines VLA with a rule-based controller can successfully perform the entire disassembly operation. These findings highlight the current limitations of VLA models in handling the dexterity and precision required for robotic EoL product disassembly. By offering a detailed analysis of the observed results, this study provides insights that may inform future research to address current challenges and advance end-to-end robotic automated disassembly.

机器人拆解VLA模型端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。