构建真实声学场景下的助听器数据集,用于算法评估与训练。
HIDVAS: A Hearing Instrument Dataset in Various Acoustical Scenarios for Algorithm Evaluation and Training

- 用8个外置扬声器、2个麦克风和假人头模拟真实听觉环境。
- 覆盖4种耳塞+3种混响条件+2个房间,共生成192组音频数据。
- 适合助听器算法研发者、语音处理研究人员使用。
为评估音频信号处理算法并训练数据驱动模型(如助听器应用),可使用仿真或实录数据。尽管仿真数据可通过数学模型大规模生成,但实录数据更能反映真实场景。本文介绍听力仪器多声学场景数据集(HIDVAS),包含8个外置扬声器、2个外置麦克风及一个假人头,假人头上安装了双麦克风的耳后式(BTE)助听器壳体,耳道内插入受话器式(RIC)助听器发声单元,且在鼓膜位置设有麦克风。通过扫频正弦记录计算每对麦克风-扬声器的脉冲响应,并播放男/女语音、语音形状噪声、歌唱声、弦乐、管乐、打击乐等音频源,同步采集所有麦克风信号。实验在一种房间(混响时间T30=0.09s, 0.47s, 0.73s)中进行4种耳塞(开放、半开放、封闭、无RIC),并在另一房间(T30=1.48s)中重复一次。共生成192组数据。以“助听器在盒中”为例,展示三个典型应用场景。
原文摘要 · Abstract (English)
To evaluate the performance of audio signal processing algorithms and to train data-driven algorithms, e.g., as applied in hearing instruments, either simulated or recorded data can be used. While large batches of simulated data can be generated using mathematical models, recorded data provide a more adequate representation of real-life scenarios. Therefore, in this paper, the Hearing Instrument Dataset in Various Acoustical Scenarios (HIDVAS) is introduced. This dataset consists of both impulse responses and audio recordings using eight external loudspeakers, two external microphones, and a dummy head. On this dummy head behind-the-ear (BTE) hearing instrument shells with two microphones per shell are mounted, and in the dummy head's ears receiver-in-canal (RIC) hearing instrument loudspeakers are inserted. The dummy head also contains microphones located at its eardrum. The impulse responses have been computed from a swept-sine recording for each microphone-loudspeaker pair, and the audio recordings have been obtained by playing back audio (male and female speech, speech shaped noise, singing voice, stringed instrument, wind instrument, and percussion instrument) through each individual loudspeaker and recording simultaneously using all microphones. These recordings have been repeated for four hearing instrument domes (open, semi-open, closed, and no-RIC) in three reverberation conditions in one room (T30 = 0.09 s, T30 = 0.47 s, and T30 = 0.73 s), and in one reverberation condition in a different room (T30 = 1.48 s). The usage of the dataset as a `hearing instrument in a box' is exemplified with three example use cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。