arXiv:2602.13440cs.CVcs.RO2026-02中稿 · European Robotics …被引 1

为室内无人机设计实时学习新物体的方法,避免遗忘。

Learning on the Fly: Replay-Based Continual Object Perception for Indoor Drones

  • 用重放策略在有限内存下持续学习新物体
  • 5%内存重放时准确率达82.96%
  • 适合资源受限的无人机实时视觉系统

自主智能体如室内无人机需在实时中学习新物体类别,同时限制灾难性遗忘,这推动了类增量学习(CIL)的发展。然而,多数无人飞行器(UAV)数据集聚焦于户外场景,缺乏时间连贯的室内视频。本文构建了一个包含14,400帧的室内数据集,涵盖无人机与地面车辆的影像,通过半自动标注流程实现98.6%的首轮标注一致率,再经人工最终校验。基于此数据集,我们评估了三种基于重放的CIL策略:经验重放(ER)、最大干扰检索(MIR)和遗忘感知重放(FAR),采用YOLOv11-nano作为资源高效检测器,适用于部署受限的无人机平台。在严格内存预算(5-10%重放)下,FAR表现最优,5%重放时平均准确率(ACC,跨增量mAP_{50-95})达82.96%。梯度加权类激活映射(Grad-CAM)分析显示,在混合场景中注意力随类别转移,导致无人机定位质量下降。实验进一步表明,基于重放的持续学习可有效应用于边缘空基系统。本工作贡献了一个保留时间连贯性的室内无人机视频数据集,并在有限重放预算下评估了重放式持续学习方法。

原文摘要 · Abstract (English)

Autonomous agents such as indoor drones must learn new object classes in real-time while limiting catastrophic forgetting, motivating Class-Incremental Learning (CIL). However, most unmanned aerial vehicle (UAV) datasets focus on outdoor scenes and offer limited temporally coherent indoor videos. We introduce an indoor dataset of $14,400$ frames capturing inter-drone and ground vehicle footage, annotated via a semi-automatic workflow with a $98.6\%$ first-pass labeling agreement before final manual verification. Using this dataset, we benchmark 3 replay-based CIL strategies: Experience Replay (ER), Maximally Interfered Retrieval (MIR), and Forgetting-Aware Replay (FAR), using YOLOv11-nano as a resource-efficient detector for deployment-constrained UAV platforms. Under tight memory budgets ($5-10\%$ replay), FAR performs better than the rest, achieving an average accuracy (ACC, $mAP_{50-95}$ across increments) of $82.96\%$ with $5\%$ replay. Gradient-weighted class activation mapping (Grad-CAM) analysis shows attention shifts across classes in mixed scenes, which is associated with reduced localization quality for drones. The experiments further demonstrate that replay-based continual learning can be effectively applied to edge aerial systems. Overall, this work contributes an indoor UAV video dataset with preserved temporal coherence and an evaluation of replay-based CIL under limited replay budgets. Project page: https://spacetime-vision-robotics-laboratory.github.io/learning-on-the-fly-cl

持续学习无人机目标检测重放策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。