arXiv:2506.05163cs.CV2025-06被引 18

构建首个多模态无人机感知数据集,支持高速运动与复杂光照下的检测跟踪。

FRED: The Florence RGB-Event Drone Dataset

  • 融合RGB视频与事件流,实现高时序分辨率的多模态感知。
  • 包含7小时以上标注轨迹,覆盖5种无人机型号及雨天等挑战场景。
  • 提供标准评估协议,助力高速无人机感知研究与可复现对比。

小型、快速且轻量化的无人机对传统RGB相机构成挑战,尤其在快速运动和复杂光照条件下表现不佳。事件相机具备高时间分辨率和动态范围,是理想解决方案,但现有基准数据集往往缺乏精细的时间分辨率或无人机特有运动模式,制约了相关进展。本文提出佛罗伦萨RGB-事件无人机数据集(FRED),一个专为无人机检测、跟踪与轨迹预测设计的新型多模态数据集,融合RGB视频与事件流。FRED包含超过7小时密集标注的无人机轨迹,使用5种不同无人机模型,并涵盖雨天及恶劣光照等挑战性场景。我们提供了各任务的标准评估协议与度量指标,支持可复现的基准测试。作者期望FRED能推动高速无人机感知与多模态时空理解的研究发展。

原文摘要 · Abstract (English)

Small, fast, and lightweight drones present significant challenges for traditional RGB cameras due to their limitations in capturing fast-moving objects, especially under challenging lighting conditions. Event cameras offer an ideal solution, providing high temporal definition and dynamic range, yet existing benchmarks often lack fine temporal resolution or drone-specific motion patterns, hindering progress in these areas. This paper introduces the Florence RGB-Event Drone dataset (FRED), a novel multimodal dataset specifically designed for drone detection, tracking, and trajectory forecasting, combining RGB video and event streams. FRED features more than 7 hours of densely annotated drone trajectories, using 5 different drone models and including challenging scenarios such as rain and adverse lighting conditions. We provide detailed evaluation protocols and standard metrics for each task, facilitating reproducible benchmarking. The authors hope FRED will advance research in high-speed drone perception and multimodal spatiotemporal understanding.

多模态无人机感知事件相机数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。