构建首个真实场景无人机音频数据集,助力复杂环境下人类存在检测。
DroneAudioset: An Audio Dataset for Drone-based Search and Rescue
- 采集23.5小时真实无人机音频,覆盖多种机型、音量与环境条件
- 信号噪声比低至-57.2 dB,模拟极端干扰下的语音识别挑战
- 支持降噪算法开发,适合无人机搜救与音频系统设计研究者
无人机在搜救任务中日益重要,但视觉方法在低可见度或遮挡下易失效。音频感知虽有潜力,却受严重自身噪声干扰。现有数据集要么多样性不足,要么为合成数据,缺乏真实声学交互,且无统一采集标准。为此,我们发布DroneAudioset,一个涵盖23.5小时标注音频的综合性无人机音频数据集,信号噪声比范围从-57.2 dB到-2.5 dB,覆盖多种无人机型号、油门设置、麦克风配置及环境。该数据集支持在恶劣条件下进行噪声抑制与人类存在分类方法的开发与系统评估,同时可指导麦克风布局等实际系统设计,推动无人机听觉系统的发展。数据集已开源(MIT许可)。
原文摘要 · Abstract (English)
Unmanned Aerial Vehicles (UAVs) or drones, are increasingly used in search and rescue missions to detect human presence. Existing systems primarily leverage vision-based methods which are prone to fail under low-visibility or occlusion. Drone-based audio perception offers promise but suffers from extreme ego-noise that masks sounds indicating human presence. Existing datasets are either limited in diversity or synthetic, lacking real acoustic interactions, and there are no standardized setups for drone audition. To this end, we present DroneAudioset (The dataset is publicly available at https://huggingface.co/datasets/ahlab-drone-project/DroneAudioSet/ under the MIT license), a comprehensive drone audition dataset featuring 23.5 hours of annotated recordings, covering a wide range of signal-to-noise ratios (SNRs) from -57.2 dB to -2.5 dB, across various drone types, throttles, microphone configurations as well as environments. The dataset enables development and systematic evaluation of noise suppression and classification methods for human-presence detection under challenging conditions, while also informing practical design considerations for drone audition systems, such as microphone placement trade-offs, and development of drone noise-aware audio processing. This dataset is an important step towards enabling design and deployment of drone-audition systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。