用DINOv3提升蓝莓采摘机器人视觉感知,验证其作为语义骨干的优劣
DINOv3 Visual Representations for Blueberry Perception Toward Robotic Harvesting
- 以冻结的DINOv3为骨干,配合轻量解码器统一评估多种蓝莓视觉任务
- 分割任务性能随骨干规模提升且稳定,检测受限于目标尺度与定位兼容性
- 适合关注农业视觉任务中语义特征与空间建模匹配的研究者
通过大规模自监督学习训练的视觉基础模型在视觉感知中展现出强泛化能力,但在农业场景中的实际作用与性能边界仍不明确。本文评估了DINOv3作为冻结骨干在蓝莓机器人采摘相关视觉任务中的表现,包括果实与碰伤分割、果实与果簇检测。在统一协议下使用轻量解码器,分割任务得益于稳定的局部特征表示,性能随骨干规模提升;而检测受目标尺度变化、补丁离散化及定位兼容性限制。果簇检测失败暴露了对空间聚合关系建模的局限性。整体而言,DINOv3不应被视为端到端任务模型,而应作为依赖下游空间建模的语义骨干,其效果取决于与果实尺度及聚合结构一致的空间设计,为蓝莓机器人采摘提供指导。代码与数据集将在接受后公开。
原文摘要 · Abstract (English)
Vision Foundation Models trained via large-scale self-supervised learning have demonstrated strong generalization in visual perception; however, their practical role and performance limits in agricultural settings remain insufficiently understood. This work evaluates DINOv3 as a frozen backbone for blueberry robotic harvesting-related visual tasks, including fruit and bruise segmentation, as well as fruit and cluster detection. Under a unified protocol with lightweight decoders, segmentation benefits consistently from stable patch-level representations and scales with backbone size. In contrast, detection is constrained by target scale variation, patch discretization, and localization compatibility. The failure of cluster detection highlights limitations in modeling relational targets defined by spatial aggregation. Overall, DINOv3 is best viewed not as an end-to-end task model, but as a semantic backbone whose effectiveness depends on downstream spatial modeling aligned with fruit-scale and aggregation structures, providing guidance for blueberry robotic harvesting. Code and dataset will be available upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。