用视觉语言动作策略实现温室草莓采摘,仅需4小时真实数据即达74%成功率。
HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild
- 三视角视觉闭环控制,无需深度图与标定
- 全微调后成功率74.0%,单次采摘32.6秒,损伤率4.1%
- 适合机器人采摘、开放环境部署的实证研究
本工作首次将视觉-语言-动作(VLA)策略迁移至真实温室桌面草莓采摘任务,该任务具有长时程、无结构特性,受遮挡和镜面反射挑战。我们在HarvestFlex平台上构建了端到端闭环系统,采用三视图RGB感知(两个固定场景视角+腕装视角),刻意避免使用深度云与显式几何标定。收集了3.71小时的VR遥操作示范数据(227个任务周期),对pi_0、pi_0.5及WALL-OSS模型进行全微调和LoRA微调。在统一的50次真实温室测试协议下,pi_0.5全微调版本取得74.0%的成功率,单次采摘耗时32.6秒,损伤率为4.1%。异步推理-控制解耦进一步提升性能。结果表明,仅用不到四小时真实数据即可实现非平凡的闭环采摘,但仍受限于近距离观测能力下降和接触动力学不匹配问题。演示视频见:https://youtu.be/bN8ZowZKPMI。
原文摘要 · Abstract (English)
This work presents the first study on transferring vision-language-action (VLA) policies to real greenhouse tabletop strawberry harvesting, a long-horizon, unstructured task challenged by occlusion and specular reflections. We built an end-to-end closed-loop system on the HarvestFlex platform using three-view RGB sensing (two fixed scene views plus a wrist-mounted view) and intentionally avoided depth clouds and explicit geometric calibration. We collected 3.71 h of VR teleoperated demonstrations (227 episodes) and fine-tuned pi_0, pi_0.5, and WALL-OSS with full fine-tuning and LoRA. Under a unified 50 trials real-greenhouse protocol and metrics spanning completion, pi_0.5 with full fine-tuning achieved success rate of 74.0% with 32.6 s/pick and damage rate of 4.1%. Asynchronous inference-control decoupling further improved performance over synchronous deployment. Results showed non-trivial closed-loop picking with fewer than four hours of real data, while remaining limited by close-range observability loss and contact-dynamics mismatch. A demonstration video is available at: https://youtu.be/bN8ZowZKPMI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。