通过区分任务相关与无关区域,提升农业机器人视觉模仿学习的泛化能力
Task-Relevant and Irrelevant Region-Aware Augmentation for Generalizable Vision-Based Imitation Learning in Agricultural Manipulation
- 将视觉输入分为任务相关和无关区域,分别进行针对性增强
- 在未见视觉条件下,采摘成功率显著高于基线方法
- 适合需要强泛化能力的农业机器人视觉控制任务
基于视觉的模仿学习在机器人操作中展现出潜力,但在实际农业任务中泛化能力仍受限。这主要源于示范数据稀少以及由作物外观多样性与背景差异引起的显著视觉域差距。为此,我们提出双区域增强模仿学习(DRAIL),一种面向农业操作中可泛化视觉模仿学习的区域感知增强框架。DRAIL 显式将视觉观测分为任务相关与任务无关区域:任务相关区域在领域知识驱动下进行增强以保留关键视觉特征,而任务无关区域则被大幅随机化以抑制虚假背景关联。通过联合处理这两类视觉变化,DRAIL 促使策略学习依赖于任务本质特征而非偶然视觉线索。我们在基于扩散策略的视觉-运动控制器上通过机器人实验评估 DRAIL,在人工蔬菜采摘和真实生菜缺陷叶摘除准备任务中均表现出在未见视觉条件下的成功率达显著提升。进一步的注意力分析与表示泛化指标表明,所学策略更依赖任务本质视觉特征,从而增强了鲁棒性与泛化性能。
原文摘要 · Abstract (English)
Vision-based imitation learning has shown promise for robotic manipulation; however, its generalization remains limited in practical agricultural tasks. This limitation stems from scarce demonstration data and substantial visual domain gaps caused by i) crop-specific appearance diversity and ii) background variations. To address this limitation, we propose Dual-Region Augmentation for Imitation Learning (DRAIL), a region-aware augmentation framework designed for generalizable vision-based imitation learning in agricultural manipulation. DRAIL explicitly separates visual observations into task-relevant and task-irrelevant regions. The task-relevant region is augmented in a domain-knowledge-driven manner to preserve essential visual characteristics, while the task-irrelevant region is aggressively randomized to suppress spurious background correlations. By jointly handling both sources of visual variation, DRAIL promotes learning policies that rely on task-essential features rather than incidental visual cues. We evaluate DRAIL on diffusion policy-based visuomotor controllers through robot experiments on artificial vegetable harvesting and real lettuce defective leaf picking preparation tasks. The results show consistent improvements in success rates under unseen visual conditions compared to baseline methods. Further attention analysis and representation generalization metrics indicate that the learned policies rely more on task-essential visual features, resulting in enhanced robustness and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。