实测工业包装任务中视觉语言动作模型的部署难题与优化路径
A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons

- 通过迭代数据收集与微调,将预训练模型适配到工厂真实场景
- 累计2535次实验(10小时)后发现反复出现的失败模式
- 适合关注机器人部署落地、工业自动化应用的研究者
视觉-语言-动作(VLA)策略展现出强大的操作能力,但其在实际部署中的可靠性常受制约。本文报告了在西门子埃尔兰根工厂(GWE)进行的工业包装任务部署案例,机器人需从杂乱堆叠中取出透明配件袋,插入纸箱残余腔体,并确保袋及内容物低于闭合平面。目标是通过迭代微调与部署驱动的优化,将预训练的Pi0.5策略适配至单一产线任务。整个流程包含数据采集、清洗、微调、评估与针对性补救数据收集的循环。共积累2535个任务样本(总计10小时),本文提供了一项关于工厂级VLA部署的实证分析,揭示了反复出现的失败模式,并总结出改进部署流程的关键经验。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) policies have shown promising manipulation capabilities, yet their practical impact is often limited by the reliability demands of real-world deployment. We present a deployment study of an industrial packaging task at Siemens Factory (GWE, Erlangen, Germany), where a robot must pick a transparent accessory bag from a cluttered pile, insert it into the remaining cavity of a cardboard package, and ensure that the bag and its contents remain below the closing plane. Our goal is to understand the practical effort required to adapt a pretrained Pi0.5 policy to a single factory-floor task through iterative fine-tuning and deployment-driven refinement. The pipeline consists of repeated loops of data collection, curation, fine-tuning, evaluation, and targeted recovery data collection. We have accumulated 2535 episodes (10 hours) from the on-site factory settings. In this paper, we contribute an empirical account of a factory-floor VLA deployment, highlighting recurring failure modes and lessons that inform how to improve the deployment workflow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。