用视觉语言动作模型实现真实尺寸双臂家具组装,突破长时序控制难题。
FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model

- 构建视觉语言动作模型,联合预测动作与进度信号,自动切换组装子任务。
- 仿真成功率从48%提升至80%,真实机器人上最难任务仅下降16%。
- 首个真实尺度双臂组装系统,适合工业自动化与机器人学习研究者。
当前机器人家具组装研究多集中于玩具级场景或单臂操作。本文提出FurnitureVLA,首个基于视觉语言动作模型(VLAs)的真实尺度双臂家具组装系统。我们形式化该任务,开发可扩展的仿真数据生成与评估流水线,并构建虚拟现实遥操作平台,由单人完成双臂控制,采集高质量真实示范数据。针对长达7个子任务、1550步控制的超长时序组装问题,提出增强进度感知的VLA模型,在语义化子任务上微调,联合预测动作与连续进度信号,实现自动子任务切换并减少推理中误差累积。进一步研究感知与控制设计因素对真实尺度组装精度的影响。FurnitureVLA在三种家具类型上使平均仿真成功率从48%提升至80%,额外通过设计优化获得21%提升;在真实Kinova Gen3平台上验证,最困难任务仅下降16%。
原文摘要 · Abstract (English)
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the first systematic study of real-scale bimanual furniture assembly using Vision-Language-Action models (VLAs). We formalize the task, develop a scalable simulation pipeline for expert data generation and evaluation, and build a VR teleoperation system for single-operator bimanual control to collect high-quality real-world demonstrations. To address extreme long-horizon assembly with up to 7 subtasks and 1550 control steps, we propose a progress-enhanced VLA, finetuned on semantically grounded subtasks, that jointly predicts actions and a continuous progress signal, enabling automatic subtask transitions and reducing compounding errors during inference. We further study perception and control design factors that critically affect precision in real-scale assembly. FurnitureVLA improves average simulation success from 48% to 80% compared to baselines across three furniture types, with an additional 21% gain from our design factor study. We validate on a real Kinova Gen3 platform with only 16% drop on the hardest task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。