arXiv:2602.01389cs.RO2026-02中稿 · ICRA

用3D地图生成伪标签,让机器人自适应新环境的语义分割。

Instance-Guided Unsupervised Domain Adaptation for Robotic Semantic Segmentation

  • 基于3D地图生成多视角一致的伪标签
  • 用基础模型修复实例级不一致性,提升标签质量
  • 无需真实标签,适合长期运行的机器人部署

语义分割网络在机器人感知中至关重要,但当部署环境的视觉分布与训练数据不同时,性能会下降。无监督域自适应(UDA)通过利用机器人长期运行中自然收集的大规模数据,在无外部标注的情况下适应目标环境。现有方法依赖环境地图的多视角一致性来无监督微调模型,缓解域偏移问题。然而,这些方法对跨视角实例级不一致仍敏感。本文提出新方法:从体素化3D地图生成多视角一致的伪标签,并利用基础模型的零样本实例分割能力进行标签精炼,强化实例级一致性。精炼后的标注作为自监督微调的监督信号,使机器人可在部署时自适应其感知系统。在真实世界数据上的实验表明,本方法持续优于基于多视角一致性的先进UDA基线,且无需目标域任何真值标签。

原文摘要 · Abstract (English)

Semantic segmentation networks, which are essential for robotic perception, often suffer from performance degradation when the visual distribution of the deployment environment differs from that of the source dataset on which they were trained. Unsupervised Domain Adaptation (UDA) addresses this challenge by adapting the network to the robot's target environment without external supervision, leveraging the large amounts of data a robot might naturally collect during long-term operation. In such settings, UDA methods can exploit multi-view consistency across the environment's map to fine-tune the model in an unsupervised fashion and mitigate domain shift. However, these approaches remain sensitive to cross-view instance-level inconsistencies. In this work, we propose a method that starts from a volumetric 3D map to generate multi-view consistent pseudo-labels. We then refine these labels using the zero-shot instance segmentation capabilities of a foundation model, enforcing instance-level coherence. The refined annotations serve as supervision for self-supervised fine-tuning, enabling the robot to adapt its perception system at deployment time. Experiments on real-world data demonstrate that our approach consistently improves performance over state-of-the-art UDA baselines based on multi-view consistency, without requiring any ground-truth labels in the target domain.

语义分割域自适应机器人感知自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。