让音乐采样按需换音色,保持原曲结构与音质
AdaTT: Text-Guided Instrument Timbre Transfer with Target-Adaptive Structural Control

- 根据目标乐器特性动态调节音高和响度控制强度
- 在多种音色迁移场景中实现高保真度与自然感
- 适合音乐生成、音色编辑等需要精准控制的场景
本文针对细粒度结构条件下乐器音色迁移中的音色模糊问题。我们认为,该问题源于特定乐器在这些结构条件下的表达细节与目标音色属性之间的冲突。例如,将小提琴以音高为主导的颤音轮廓施加于本应以响度为主导颤音的长笛,会损害音色保真度。为此,我们提出AdaTT,一种基于ControlNet框架的目标自适应系统,通过文本提示选择性地调整帧级音高与响度控制的影响权重,以匹配目标乐器身份。同时,我们设计了一套半自动数据构建流程,指导模型识别需变换或保留的表达细节。实验表明,AdaTT在保持乐谱级内容的同时,显著提升了音色保真度与自然度。音频样本见https://dabinkim0.github.io/adatt/。
原文摘要 · Abstract (English)
This paper addresses timbral ambiguity in instrument timbre transfer under fine-grained structural conditions. We argue this issue stems from instrument-specific expressive details in these conditions, which conflict with the target timbral properties. For example, imposing a violin's pitch-dominant vibrato contours onto a flute, which naturally exhibits loudness-dominant vibrato, impairs timbral fidelity. We propose AdaTT, a target-adaptive system that ensures high timbral fidelity across diverse timbre transfer scenarios within the ControlNet scheme. It selectively scales the frame-wise influence of pitch and loudness controls via text prompts to match the target instrument's identity. We also present a semi-automatic data construction pipeline to teach the model which expressive details to transform or preserve. Results show AdaTT achieves superior timbral fidelity and naturalness while retaining score-level content. Audio samples are available at https://dabinkim0.github.io/adatt/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。