首个标注真实场景噪声事件的多语言语音数据集,助力噪声鲁棒语音识别。
VAANI Noise Event Dataset: A curated spontaneous speech dataset annotated with timestamps for noise events
- 基于165个印度地区的实地录音,构建多语言真实语音与环境噪声同步数据
- 细粒度标注7类重叠噪声事件的时间戳,涵盖动物、交通、婴儿等场景
- 适用于噪声鲁棒ASR、声学事件检测等任务,适合多语言语音研究者
现有公开声学事件数据集大多针对通用音频分类或纯净语音分离,较少提供在真实自发语音上叠加的精确时间戳噪声标注。本文提出VAANI噪声事件时间戳数据集,基于项目VAANI在105种印地语系语言中覆盖165个印度地区的实地录音构建。与合成混合数据集不同,VAANI同步记录真实语音与环境噪声,并为每段录音标注细粒度的噪声起止时间戳,将噪声事件归入7类紧凑语义分类:动物、交通、婴儿/儿童、音乐、信号/警报、电器、非语音人声。该数据集结合了自发多语言印地语语音、真实的区域声景与可与语音目标重叠的段级噪声标签,填补了当前数据集在噪声鲁棒自动语音识别(ASR)、声学事件检测(SED)和语音增强任务中的空白。本文将VAANI与九个常用数据集(包括WHAM!、AVA-Speech、MUSAN、FSD50K、CHiME-6、AudioSet、DESED、iNoise和Kathbath-Noisy)进行对比,并详细描述了标注协议与质量控制流程。
原文摘要 · Abstract (English)
Most public sound-event corpora are optimized either for general audio tagging or for clean speech separation, and comparatively few provide strong timestamped noise annotations layered directly on top of spontaneous, real-world speech. We present the VAANI Noise Event Timestamp Dataset, a derived annotation layer built on Project VAANI field recordings of spontaneous speech collected across 165 Indian districts in 105 languages. Unlike synthetically mixed corpora, VAANI captures speech and ambient noise in situ and simultaneously, and annotates each recording with fine-grained start/end timestamps for overlapping background noise events organized into a compact seven-class semantic taxonomy: animal, traffic, baby/child, music, signal/alarm, appliance, and non-speech human. This combination of spontaneous multilingual Indic speech, authentic regional soundscapes, and span-level noise tags that may overlap with speech targets tasks that existing datasets address only partially: noise-robust Automatic Speech Recognition (ASR), sound event detection (SED), and speech enhancement. We position VAANI against nine widely used corpora and benchmarks, including WHAM!, AVA-Speech, MUSAN, FSD50K, CHiME-6, AudioSet, DESED, the India-specific iNoise noise database, and the Kathbath-Noisy noisy-ASR benchmarks, and describe the annotation protocol and quality-control procedure used to produce the timestamped tags.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。