用多样性搜索生成创新音频,让音乐创作更高效。
Quality-Diversity Search in Sound Generation: Investigating Innovation Engines for Audio Exploration

- 结合质量-多样性算法与判别模型,自动探索未知音色空间。
- 多频段专用CPPN设计使网络更简单,性能不降。
- 适合音乐人、声音设计师和生成艺术研究者使用。
本研究针对作曲家和声音设计师在实现音乐目标时面临的挑战,提出一种基于进化过程的自动化音色探索方法,通过促进多样性以激发意外发现。系统融合质量-多样性(QD)算法与监督判别模型,受创新引擎算法启发,探索不同配置及合成方式与判别模型间的相互作用。研究对比了组合模式产生网络(CPPNs)与数字信号处理(DSP)图的关系,引入多专用CPPN分频段处理的新方法,在保持单个大网络性能的同时简化结构。通过分析音乐与非音乐情境间目标切换的进化路径,揭示了谱系如何穿越低概率路径抵达当前最优。扩展先前研究的行为空间至多种声音持续时间,发现时间域内的专业化现象。结果表明,结合多维表型精英存档(MAP-Elites)与深度学习分类器的CPPN与DSP图可生成丰富多样、跨时空维度具有创新性的合成音色。生成结果通过在线探索器与音频文件展示,并在作曲实验中验证其跨时长与情境的创作潜力。
原文摘要 · Abstract (English)
This study addresses the challenges composers and sound designers face in creating and refining tools to achieve their musical goals. Using evolutionary processes to promote diversity and foster serendipitous discoveries, we automate the search through uncharted sonic spaces for sound discovery, arguing that diversity-promoting algorithms can bridge the gap between the theoretical realisation and practical accessibility of sounds. We describe a system for generative sound synthesis combining Quality Diversity (QD) algorithms with a supervised discriminative model, inspired by the Innovation Engine algorithm, and explore different configurations and the interplay between the chosen synthesis approach and the discriminative model. We examine the interaction between Compositional Pattern Producing Networks (CPPNs) and Digital Signal Processing (DSP) graphs, introducing a novel approach that uses multiple specialised CPPNs for different frequency ranges; this yields simpler networks while maintaining performance comparable to single-CPPN setups. We also investigate evolutionary stepping stones by analysing goal switches between musical and non-musical contexts, revealing how lineages traverse unlikely paths to current elites. Expanding the behaviour space of a previous study to include various sound durations, we uncover specialisation within temporal niches. Results indicate that CPPN and DSP graphs coupled with a Multi-dimensional Archive of Phenotypic Elites (MAP-Elites) and a deep learning classifier can generate a substantial variety of synthetic sounds, diverse and innovative across temporal and contextual dimensions. We present the generated sound objects through an online explorer and as rendered sound files, and, in the context of music composition, an experimental application that showcases their creative potential across various durations and contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。