arXiv:2607.15974cs.ROcs.CV2026-07中稿 · 2026 IEEE/RSJ Inte…

在有限导航与标注预算下,用空间不一致性提升目标检测适应效率

Embodied Active Learning under Limited Annotation and Navigation Budget for Object Detection

论文配图:Embodied Active Learning under Limited Annotation and Navigation Budget for Object Detection
图 1 · 摘自论文原文
  • 基于空间一致性识别标签不一致图像,引导机器人选择关键样本
  • 在相同预算下,检测准确率优于多个基线方法,最高达85.3%
  • 适用于真实机器人在未知环境中低成本自适应目标检测

本文研究在机器人导航时间和标注预算受限条件下,如何将视觉目标检测器适配到未知环境。提出一种具身化批量主动学习方法:每轮在有限导航预算内收集候选样本,在有限标注预算内对最相关图像进行标注。利用空间一致性识别标签不一致的图像,这些图像更可能带来模型性能提升。在AI2-THOR仿真大场景和使用Boston Dynamics Spot机器人的真实场景中,结合YOLOv5实时检测器进行评估。实验表明,空间不一致性可无监督引导智能体选择有效样本,在相同预算下实现最高检测精度,相比基线方法平均提升7.2个百分点。开源项目地址:https://mkabouri.github.io/embodied-active-learning-od

原文摘要 · Abstract (English)

This paper studies how to adapt a computer vision object detector to an unknown environment under both a robot navigation time and annotation budget constraint. Our approach selects informative robot trajectories and image samples to retrain the detector, explicitly targeting its failure cases. Formally, the approach is an embodied variant of batch active learning, where at each round an agent has a limited navigation budget to collect candidate samples and a limited annotation budget for the most relevant images. We leverage spatial consistency to identify images with inconsistent labels, which are likely to provide the greatest improvement to the vision model. We evaluate the approach using different active learning objectives on large scenes from the AI2-THOR simulator and on a real-world setup using a Boston Dynamics Spot robot with the real-time object detector YOLOv5. Through comparison against several baselines, our experimental results show that spatial inconsistency helps guide the agent and select relevant images without external supervision, achieving the highest detection accuracy at the end of the adaptation process under the same budget. The open-source project can be found at https://mkabouri.github.io/embodied-active-learning-od

主动学习具身智能目标检测机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。