arXiv:2410.01481cs.SDcs.AI2024-10ICLR被引 13

SonicSim生成可定制的移动声源语音数据,提升模型真实场景泛化能力。

SonicSim: A customizable simulation platform for speech processing in moving sound source scenarios

  • 基于Habitat-sim构建,支持场景/麦克风/声源多层级自定义调整
  • 在5小时真实数据对比下,SonicSet训练模型更接近真实表现
  • 适合语音分离增强模型在动态声源下的训练与评估

在移动声源条件下系统评估语音分离与增强模型需要大量多样数据。然而真实数据集常不足,合成数据虽规模大但缺乏声学真实感,均难满足实际需求。为此,我们提出SonicSim,一个基于具身智能仿真平台Habitat-sim的合成工具包,用于生成高度可定制的移动声源数据。SonicSim支持场景级、麦克风级和声源级多层级调节,实现数据多样性提升。基于此,我们构建了基准数据集SonicSet,融合LibriSpeech、FSD50K、FMA及Matterport3D中的90个场景。为分析合成与真实数据差异,我们从SonicSet验证集选取5小时原始非混响数据,并录制真实世界语音分离数据集,作为对比参考。针对语音增强,我们使用真实数据集RealMAN验证SonicSet与其他合成数据集间的声学差距。结果表明,基于SonicSet训练的模型在真实场景中泛化性能优于其他合成数据集。代码已开源:https://cslikai.cn/SonicSim/

原文摘要 · Abstract (English)

Systematic evaluation of speech separation and enhancement models under moving sound source conditions requires extensive and diverse data. However, real-world datasets often lack sufficient data for training and evaluation, and synthetic datasets, while larger, lack acoustic realism. Consequently, neither effectively meets practical needs. To address this issue, we introduce SonicSim, a synthetic toolkit based on the embodied AI simulation platform Habitat-sim, designed to generate highly customizable data for moving sound sources. SonicSim supports multi-level adjustments, including scene-level, microphone-level, and source-level adjustments, enabling the creation of more diverse synthetic data. Leveraging SonicSim, we constructed a benchmark dataset called SonicSet, utilizing LibriSpeech, Freesound Dataset 50k (FSD50K), Free Music Archive (FMA), and 90 scenes from Matterport3D to evaluate speech separation and enhancement models. Additionally, to investigate the differences between synthetic and real-world data, we selected 5 hours of raw, non-reverberant data from the SonicSet validation set and recorded a real-world speech separation dataset, providing a reference for comparing SonicSet with other synthetic datasets. For speech enhancement, we utilized the real-world dataset RealMAN to validate the acoustic gap between SonicSet and existing synthetic datasets. The results indicate that models trained on SonicSet generalize better to real-world scenarios compared to other synthetic datasets. The code is publicly available at https://cslikai.cn/SonicSim/.

语音分离仿真平台数据生成具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。