arXiv:2507.19202cs.SDcs.LG2025-07中稿 · ISMIR 2025 Late Br…被引 5

用隐向量重构颗粒合成,让声音保结构又换音色。

Latent Granular Resynthesis using Neural Audio Codecs

  • 在隐向量层面重构颗粒合成,用编码后的声学片段匹配目标音频
  • 无需训练即可实现音色迁移,保留目标时序结构并避免拼接断点
  • 适合音乐创作与声音设计,支持任意音频素材实时实验

我们提出一种新型创意音频重合成技术,将颗粒合成的概念延伸至隐向量层面。通过将源音频语料库编码为隐向量片段,构建一个“颗粒码本”;再将目标音频的每个隐向量颗粒与码本中最接近的条目匹配,生成混合序列后解码输出。新方法无需模型训练,适用于多样音频材料,且因编码器在解码时的隐式插值,自然避免了传统拼接合成中的不连贯问题。补充材料及概念验证实现已公开于 https://github.com/naotokui/latentgranular/ ,用户可通过 Hugging Face 空间在线实验自定义声音:https://huggingface.co/spaces/naotokui/latentgranular。

原文摘要 · Abstract (English)

We introduce a novel technique for creative audio resynthesis that operates by reworking the concept of granular synthesis at the latent vector level. Our approach creates a "granular codebook" by encoding a source audio corpus into latent vector segments, then matches each latent grain of a target audio signal to its closest counterpart in the codebook. The resulting hybrid sequence is decoded to produce audio that preserves the target's temporal structure while adopting the source's timbral characteristics. This technique requires no model training, works with diverse audio materials, and naturally avoids the discontinuities typical of traditional concatenative synthesis through the codec's implicit interpolation during decoding. We include supplementary material at https://github.com/naotokui/latentgranular/ , as well as a proof-of-concept implementation to allow users to experiment with their own sounds at https://huggingface.co/spaces/naotokui/latentgranular .

音频合成隐向量颗粒合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。