无需源数据,让机器人在室内动态环境中实现高效目标检测。
Embodied Domain Adaptation for Object Detection
- 通过时序聚类优化伪标签,提升无监督域适应精度。
- 在光照、布局变化下零样本检测性能显著提升。
- 适合部署于家庭与实验室等复杂动态场景的移动机器人。
移动机器人依赖目标检测进行室内环境感知与定位,但传统封闭集方法难以应对真实家居和实验室中多样的物体及动态条件。开放词汇目标检测(OVOD)虽借助视觉语言模型扩展了标签能力,但仍面临室内环境中的域偏移问题。本文提出一种无需源数据的自监督域适应(SFDA)方法,通过时序聚类优化伪标签,采用多尺度阈值融合,并结合对比学习的均值教师框架。我们构建了面向移动机器人目标检测的领域自适应基准(EDAOD),评估在光照、布局与物体多样性连续变化下的适应能力。实验表明,该方法在零样本检测性能上取得显著提升,具备灵活应对动态室内环境的能力。
原文摘要 · Abstract (English)
Mobile robots rely on object detectors for perception and object localization in indoor environments. However, standard closed-set methods struggle to handle the diverse objects and dynamic conditions encountered in real homes and labs. Open-vocabulary object detection (OVOD), driven by Vision Language Models (VLMs), extends beyond fixed labels but still struggles with domain shifts in indoor environments. We introduce a Source-Free Domain Adaptation (SFDA) approach that adapts a pre-trained model without accessing source data. We refine pseudo labels via temporal clustering, employ multi-scale threshold fusion, and apply a Mean Teacher framework with contrastive learning. Our Embodied Domain Adaptation for Object Detection (EDAOD) benchmark evaluates adaptation under sequential changes in lighting, layout, and object diversity. Our experiments show significant gains in zero-shot detection performance and flexible adaptation to dynamic indoor conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。