arXiv:2603.06927cs.RO2026-03

用少量标注数据实现室内机器人行进路径精准分割,提升对细小障碍物的识别能力。

A Contrastive Fewshot RGBD Traversability Segmentation Framework for Indoor Robotic Navigation

  • 融合RGB图像与稀疏深度信息,通过负样本对比学习增强对障碍物的区分能力
  • 在1-shot和5-shot设置下,mIoU最高提升9%,显著优于现有方法
  • 适合需要低标注成本、高安全性的室内机器人导航场景

室内可通行性分割旨在为自主代理识别安全可通行空间,对机器人导航至关重要。纯视觉模型常无法检测如椅腿等细小障碍物,带来安全隐患。本文提出一种多模态分割框架,结合RGB图像与稀疏1D激光深度信息,捕捉几何交互以提升对挑战性障碍物的检测能力。为减少对大规模标注数据的依赖,采用少样本分割(FSS)范式,使模型能从有限标注样本中泛化。传统FSS方法仅关注正样本原型,易过拟合且泛化差。为此,我们引入负样本对比学习(NCL)分支,利用负原型(障碍物)优化自由空间预测。此外,设计两阶段注意力深度模块,实现1D深度向量与RGB图像在水平和垂直方向的对齐。在自建的室内RGB-D可通行性数据集上的大量实验表明,该方法在1-shot和5-shot设置下均超越现有SOTA FSS与RGB-D分割基线,最高提升9%的mIoU,验证了负原型与稀疏深度信息在鲁棒高效分割中的有效性。

原文摘要 · Abstract (English)

Indoor traversability segmentation aims to identify safe, navigable free space for autonomous agents, which is critical for robotic navigation. Pure vision-based models often fail to detect thin obstacles, such as chair legs, which can pose serious safety risks. We propose a multi-modal segmentation framework that leverages RGB images and sparse 1D laser depth information to capture geometric interactions and improve the detection of challenging obstacles. To reduce the reliance on large labeled datasets, we adopt the few-shot segmentation (FSS) paradigm, enabling the model to generalize from limited annotated examples. Traditional FSS methods focus solely on positive prototypes, often leading to overfitting to the support set and poor generalization. To address this, we introduce a negative contrastive learning (NCL) branch that leverages negative prototypes (obstacles) to refine free-space predictions. Additionally, we design a two-stage attention depth module to align 1D depth vectors with RGB images both horizontally and vertically. Extensive experiments on our custom-collected indoor RGB-D traversability dataset demonstrate that our method outperforms state-of-the-art FSS and RGB-D segmentation baselines, achieving up to 9\% higher mIoU under both 1-shot and 5-shot settings. These results highlight the effectiveness of leveraging negative prototypes and sparse depth for robust and efficient traversability segmentation.

少样本分割多模态感知机器人导航深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。