arXiv:2509.17684cs.CVcs.RO2025-09被引 3

用自监督视觉模型提升机器人抓取的泛化与效率

DINOv3-Diffusion Policy: Self-Supervised Large Visual Model for Visuomotor Diffusion Policy Learning

  • 用自监督DINOv3替代传统ImageNet模型做视觉编码
  • 微调后在多个任务上成功率最高提升10%
  • 无需标注数据,适合大规模机器人学习场景

本文评估了最新大尺度自监督视觉骨干DINOv3在机器人操作中的视觉-运动扩散策略学习效果。在四个基准任务(Push-T、Lift、Can、Square)上,使用统一的FiLM条件扩散策略,研究纯自监督编码器在从零训练、冻结权重和微调三种模式下是否可媲美或超越传统ImageNet预训练模型(如ResNet-18)。结果表明:(i) 微调后的DINOv3在多个任务上达到或超过ResNet-18表现;(ii) 冻结的DINOv3仍具竞争力,显示其具备强泛化先验能力;(iii) 自监督特征显著提升样本效率与鲁棒性。相比使用ResNet18,DINOv3在复杂任务Can上测试成功率最高提升10%,在Lift、PushT、Square任务上表现相当。

原文摘要 · Abstract (English)

This paper evaluates DINOv3, a recent large-scale self-supervised vision backbone, for visuomotor diffusion policy learning in robotic manipulation. We investigate whether a purely self-supervised encoder can match or surpass conventional supervised ImageNet-pretrained backbones (e.g., ResNet-18) under three regimes: training from scratch, frozen, and finetuned. Across four benchmark tasks (Push-T, Lift, Can, Square) using a unified FiLM-conditioned diffusion policy, we find that (i) finetuned DINOv3 matches or exceeds ResNet-18 on several tasks, (ii) frozen DINOv3 remains competitive, indicating strong transferable priors, and (iii) self-supervised features improve sample efficiency and robustness. These results support self-supervised large visual models as effective, generalizable perceptual front-ends for action diffusion policies, motivating further exploration of scalable label-free pretraining in robotic manipulation. Compared to using ResNet18 as a backbone, our approach with DINOv3 achieves up to a 10% absolute increase in test-time success rates on challenging tasks such as Can, and on-the-par performance in tasks like Lift, PushT, and Square.

机器人控制自监督学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。