arXiv:2509.16862cs.SDeess.AS2025-09中稿 · 2025 Asia Pacific …

将鼓声转换为人声打击乐,实现a cappella中的节奏音色还原。

Drum-to-Vocal Percussion Sound Conversion and Its Evaluation Methodology

  • 基于鼓声的音色迁移,用声学特性匹配实现人声打击乐合成。
  • 带向量量化(VQ)的RAVE模型在主观评价中更稳定一致。
  • 适用于音乐创作与人声打击乐生成,尤其适合a cappella场景。

本文定义了一项新任务:鼓声到人声打击乐(VP)的声音转换。VP通过人声模拟打击乐器,在当代无伴奏合唱音乐中常见,其声学特性(如非周期性、噪声瞬态、无语言结构)区别于言语和歌唱,因此传统语音或歌唱合成方法不适用。为此,我们将VP合成建模为从鼓声进行音色迁移的问题,利用鼓声在节奏与音色上的对应关系。为支持该方法,提出三项成功转换的核心要求:节奏保真度、音色一致性及人声打击乐自然度,并构建相应的主观评估标准。采用神经音频合成器实时音频变分自编码器(RAVE),实现两种基线转换方法,分别带有与不带向量量化(VQ)。主观实验表明,两种方法均能生成合理的人声打击乐输出,其中基于VQ的RAVE模型表现出更优的转换一致性。

原文摘要 · Abstract (English)

This paper defines the novel task of drum-to-vocal percussion (VP) sound conversion. VP imitates percussion instruments through human vocalization and is frequently employed in contemporary a cappella music. It exhibits acoustic properties distinct from speech and singing (e.g., aperiodicity, noisy transients, and the absence of linguistic structure), making conventional speech or singing synthesis methods unsuitable. We thus formulate VP synthesis as a timbre transfer problem from drum sounds, leveraging their rhythmic and timbral correspondence. To support this formulation, we define three requirements for successful conversion: rhythmic fidelity, timbral consistency, and naturalness as VP. We also propose corresponding subjective evaluation criteria. We implement two baseline conversion methods using a neural audio synthesizer, the real-time audio variational autoencoder (RAVE), with and without vector quantization (VQ). Subjective experiments show that both methods produce plausible VP outputs, with the VQ-based RAVE model yielding more consistent conversion.

声音转换人声打击乐音频合成RAVE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。