arXiv:2606.18582cs.CVcs.RO2026-06

用DINOv3+ViT-Adapter提升野外场景细粒度分割精度,获ICRA 2026冠军。

Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Leveraging DINOv3 for Robust Outdoor Scene Understanding in Field Robotics

论文配图:Technical Report for ICRA 2026 GOOSE 2D Fine-Grained Semantic Segmentation Challenge: Leveraging DINOv3 for Robust Outdoor Scene Understanding in Field Robotics
图 1 · 摘自论文原文
  • 融合DINOv3骨干网络与ViT-Adapter,辅以类别辅助损失提升特征表达
  • 多尺度翻转增强与前三名检查点集成,推理时显著提升鲁棒性
  • 在64类细粒度分类中达69.32% mIoU,适合复杂野外机器人感知

2026年国际机器人与自动化会议(ICRA)野外机器人研讨会举办的GOOSE 2D细粒度语义分割挑战赛,评估非铺装路面图像在64个细粒度类别和11个非空粗粒度类别上的密集语义分割性能。本文提出该挑战赛第一名解决方案。方法包含两项互补改进:(a) 网络层面设计,采用自监督DINOv3 ViT-L/16骨干网络、ViT-Adapter以及Mask2Former掩码分类解码器,并在全局[CLS]标记上加入粗粒度类别辅助损失;(b) 推理阶段策略,基于多尺度与水平翻转的测试时增强,结合Codabench评分选出的前三个检查点进行集成。最终在官方评测中取得76.57%的综合得分,其中细粒度类别mIoU为69.32%,粗粒度类别mIoU为83.81%,位居最终排行榜首位:www.codabench.org/competitions/14257/#/results-tab。

原文摘要 · Abstract (English)

The GOOSE 2D Fine-Grained Semantic Segmentation Challenge at the ICRA 2026 Workshop on Field Robotics evaluates dense semantic segmentation of off-road imagery over a fine-grained taxonomy of 64 classes and 11 evaluated non-void coarse categories. We present the first-place solution to this challenge. Our solution comprises two complementary improvements: (a) a network-level design that combines a self-supervised DINOv3 ViT-L/16 backbone, a ViT-Adapter, and a Mask2Former mask-classification decoder, together with a coarse-category auxiliary loss on the global [CLS] token; and (b) an inference-time aggregation strategy based on multi-scale and horizontal-flip test-time augmentation and an ensemble of the top three checkpoints selected using Codabench scores. Our method achieves an official composite score of 76.57%, consisting of 69.32% fine-class mIoU and 83.81% category-level mIoU, and ranks first on the final phase leaderboard: www.codabench.org/competitions/14257/#/results-tab.

细粒度分割DINOv3野外机器人语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。