arXiv:2509.19295eess.AScs.AI2025-09中稿 · the 10th Workshop …

在车流噪音中实现行人听觉检测,构建了超1300小时的实路数据集。

Audio-Based Pedestrian Detection in the Presence of Vehicular Noise

  • 构建1321小时路边音频数据集,含车流背景与行人标注。
  • 模型在嘈杂环境中性能下降明显,声学上下文影响显著。
  • 适合自动驾驶、智能交通领域研究者关注音频感知鲁棒性。

基于音频的行人检测是一项具有挑战性的任务,此前仅在噪声受限环境下被研究。本文提出一个新数据集,并对车载噪声环境下音频行人检测的前沿技术进行了全面分析。研究包含三项分析:(i) 在嘈杂与无噪声环境间的跨数据集评估;(ii) 噪音数据对模型性能的影响评估,强调声学上下文的作用;(iii) 模型对域外声音的预测鲁棒性评估。新数据集为综合性路边数据集,总时长达1321小时,包含丰富车流声景。每段录音均包含16kHz音频,与逐帧行人标注同步,并附带1fps视频缩略图。

原文摘要 · Abstract (English)

Audio-based pedestrian detection is a challenging task and has, thus far, only been explored in noise-limited environments. We present a new dataset, results, and a detailed analysis of the state-of-the-art in audio-based pedestrian detection in the presence of vehicular noise. In our study, we conduct three analyses: (i) cross-dataset evaluation between noisy and noise-limited environments, (ii) an assessment of the impact of noisy data on model performance, highlighting the influence of acoustic context, and (iii) an evaluation of the model's predictive robustness on out-of-domain sounds. The new dataset is a comprehensive 1321-hour roadside dataset. It incorporates traffic-rich soundscapes. Each recording includes 16kHz audio synchronized with frame-level pedestrian annotations and 1fps video thumbnails.

音频检测行人识别车载噪音数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。