arXiv:2510.04339cs.SDcs.AI2025-10中稿 · the Proceedings of…被引 1

用可交互的音色潜空间生成精准音高音乐,让创作更直观。

Pitch-Conditioned Instrument Sound Synthesis From an Interactive Timbre Latent Space

  • 分两阶段训练:先用变分自编码器学出音高与音色解耦的2维潜空间。
  • 生成模型在潜空间上控制,音高准确率高,能捕捉细微音色变化。
  • 配套网页应用可实时操控,适合音乐人和声音设计师探索新音色。

本文提出一种新型神经乐器音色合成方法,采用两阶段半监督学习框架,能够从富有表现力的音色潜空间生成音高精准、高质量的音乐样本。现有方法虽能达到制作水准,但依赖高维潜变量,难以操作且用户体验不直观。我们通过两阶段训练解决此问题:首先使用变分自编码器训练音频样本的音高-音色解耦2维表示;其次将该表示作为条件输入至基于Transformer的生成模型。所学得的2维潜空间可作为直观界面,用于探索声音景观。实验表明,该方法有效学习到解耦音色空间,实现表达性强、可控性高的音频生成,并保持高度音高准确性。用户友好性通过交互式网页应用得到验证,展示了其在下一代直观且富有创造力的音乐制作环境中的潜力:https://pgesam.faresschulz.com

原文摘要 · Abstract (English)

This paper presents a novel approach to neural instrument sound synthesis using a two-stage semi-supervised learning framework capable of generating pitch-accurate, high-quality music samples from an expressive timbre latent space. Existing approaches that achieve sufficient quality for music production often rely on high-dimensional latent representations that are difficult to navigate and provide unintuitive user experiences. We address this limitation through a two-stage training paradigm: first, we train a pitch-timbre disentangled 2D representation of audio samples using a Variational Autoencoder; second, we use this representation as conditioning input for a Transformer-based generative model. The learned 2D latent space serves as an intuitive interface for navigating and exploring the sound landscape. We demonstrate that the proposed method effectively learns a disentangled timbre space, enabling expressive and controllable audio generation with reliable pitch conditioning. Experimental results show the model's ability to capture subtle variations in timbre while maintaining a high degree of pitch accuracy. The usability of our method is demonstrated in an interactive web application, highlighting its potential as a step towards future music production environments that are both intuitive and creatively empowering: https://pgesam.faresschulz.com

音色合成潜空间交互设计音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。