arXiv:2501.17198cs.SDeess.AS2025-01被引 1

构建6000个合成音效数据集,推动程序化音频研究

6KSFx Synth Dataset

  • 基于30类声音设计6000个合成音效样本
  • 涵盖多样合成方法,支持评估框架构建
  • 适合音频生成、声学设计研究人员使用

程序化音频(又称“数字拟音”)通过计算过程从零生成声音,是音效创作的创新方法。然而,由于缺乏公开可用的数据集和模型,该技术的发展与应用受限。本文提出一个包含6000个合成音频样本的数据集,覆盖30种声音类别。数据集详细描述了各类声音所采用的合成方法,并支持构建稳健的评估体系。该资源有助于推动程序化音频的研究与开发,为研究人员、音频开发者及音效设计师提供重要支持,加速数字声音设计领域的进步。

原文摘要 · Abstract (English)

Procedural audio, often referred to as "digital Foley", generates sound from scratch using computational processes. It represents an innovative approach to sound-effects creation. However, the development and adoption of procedural audio has been constrained by a lack of publicly available datasets and models, which hinders evaluation and optimization. To address this important gap, this paper presents a dataset of 6000 synthetic audio samples specifically designed to advance research and development in sound synthesis within 30 sound categories. By offering a description of the diverse synthesis methods used in each sound category and supporting the creation of robust evaluation frameworks, this dataset not only highlights the potential of procedural audio, but also provides a resource for researchers, audio developers, and sound designers. This contribution can accelerate the progress of procedural audio, opening up new possibilities in digital sound design.

程序化音频音效生成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。