构建了高保真远场语音数据集,支持语音识别与增强研究
Treble10: A high-quality dataset for far-field speech recognition, dereverberation, and enhancement
- 用混合波-几何声学模拟10个真实房间的3000+混响响应
- 提供6类音频格式,含与LibriSpeech配对的清晰/混响语音对
- 适用于语音增强、去混响等任务,适合音频算法研发者
准确的远场语音数据集对自动语音识别(ASR)、去混响、语音增强和声源分离等任务至关重要。然而,现有数据集受限于声学真实性与可扩展性之间的权衡:实测数据虽物理真实但成本高、覆盖少,且常缺成对的清晰与混响语音;多数仿真数据依赖简化几何声学,无法再现复杂环境中的衍射、散射和干涉等关键物理现象。本文提出Treble10,一个大规模、物理精确的房间声学数据集。该数据集包含在10个全家具真实房间中模拟的超过3000个宽带房间冲激响应(RIR),采用在Treble SDK中实现的混合仿真范式,结合波基与几何声学求解器。数据集提供六类互补子集,涵盖单声道、8阶球谐编码(Ambisonics)及6通道设备级RIR,以及与LibriSpeech语句配对的预卷积混响语音场景。所有信号均以32 kHz采样,准确建模低频波效应与高频反射。Treble10弥合了实测与仿真间的现实差距,支持可复现、物理基础的评估与大规模数据增强,为远场语音任务提供基准与模板。数据集已通过Hugging Face公开发布。
原文摘要 · Abstract (English)
Accurate far-field speech datasets are critical for tasks such as automatic speech recognition (ASR), dereverberation, speech enhancement, and source separation. However, current datasets are limited by the trade-off between acoustic realism and scalability. Measured corpora provide faithful physics but are expensive, low-coverage, and rarely include paired clean and reverberant data. In contrast, most simulation-based datasets rely on simplified geometrical acoustics, thus failing to reproduce key physical phenomena like diffraction, scattering, and interference that govern sound propagation in complex environments. We introduce Treble10, a large-scale, physically accurate room-acoustic dataset. Treble10 contains over 3000 broadband room impulse responses (RIRs) simulated in 10 fully furnished real-world rooms, using a hybrid simulation paradigm implemented in the Treble SDK that combines a wave-based and geometrical acoustics solver. The dataset provides six complementary subsets, spanning mono, 8th-order Ambisonics, and 6-channel device RIRs, as well as pre-convolved reverberant speech scenes paired with LibriSpeech utterances. All signals are simulated at 32 kHz, accurately modelling low-frequency wave effects and high-frequency reflections. Treble10 bridges the realism gap between measurement and simulation, enabling reproducible, physically grounded evaluation and large-scale data augmentation for far-field speech tasks. The dataset is openly available via the Hugging Face Hub, and is intended as both a benchmark and a template for next-generation simulation-driven audio research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。