arXiv:2410.18322cs.SDcs.LG2024-10中稿 · Interspeech 2025被引 1

用频响数据统一转换麦克风音色,支持多设备互换。

Unified Microphone Conversion: Many-to-Many Device Mapping via Feature-wise Linear Modulation

  • 通过频响信息控制生成器,实现多对多设备映射。
  • 相比现有方法提升2.6%分类准确率,降低0.8%设备差异。
  • 适合需跨设备泛化的音频分类任务研究者使用。

我们提出统一麦克风转换框架,旨在增强声音事件分类(SEC)系统对设备差异的鲁棒性。此前基于CycleGAN的方法虽能模拟设备特性,但需为每对设备单独训练模型,扩展性差。本文通过将频响数据作为条件输入,结合特征逐元素线性调制(Feature-wise Linear Modulation),在无配对数据条件下实现多对多设备映射,显著提升可扩展性。此外,引入合成频响差异进一步增强实际应用适应性。实验表明,该方法在宏平均F1分数上优于当前最优方法2.6%,设备间变异性降低0.8%。

原文摘要 · Abstract (English)

We present Unified Microphone Conversion, a unified generative framework designed to bolster sound event classification (SEC) systems against device variability. While our prior CycleGAN-based methods effectively simulate device characteristics, they require separate models for each device pair, limiting scalability. Our approach overcomes this constraint by conditioning the generator on frequency response data, enabling many-to-many device mappings through unpaired training. We integrate frequency-response information via Feature-wise Linear Modulation, further enhancing scalability. Additionally, incorporating synthetic frequency response differences improves the applicability of our framework for real-world application. Experimental results show that our method outperforms the state-of-the-art by 2.6% and reduces variability by 0.8% in macro-average F1 score.

音频转换生成模型设备泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。