统一建模吉他音色迁移与还原,用变分自编码器分离音色与内容。
EG-VAE: A Unified Framework for Electric Guitar Tone Transfer and Removal

- 用变分自编码器分离录音中的帧级内容与全局音色特征
- 在客观与主观评估中均优于专用模型的音色迁移与还原效果
- 适合音乐制作人和音频工程师用于音色编辑与信号修复
电吉他音色迁移(EGTT)和音色还原(EGTR)是吉他音色建模中的两个基础任务:EGTT 将录音的音色替换为参考音色,而 EGTR 从经过处理的湿信号中恢复干信号(DI)。尽管两者密切相关,但以往研究独立处理,且效果均不理想。本文提出 EG-VAE,一个统一框架,通过变分自编码器从湿录音中解耦帧级内容与全局音色表示,实现联合建模。EGTT 通过将源内容与参考音色重新组合完成,而 EGTR 则借助一种新颖的音色掩码目标,在训练中强制内容-音色解耦,并在推理时实现音色移除。为提升对未见音色的迁移能力,第二阶段采用变分采样与音频效果增强构建平滑音色空间。客观与主观评估结果表明,EG-VAE 在音色迁移与还原任务上均优于专用基线模型。演示视频可在 https://guitar-tone-demo.vercel.app/ 查看。
原文摘要 · Abstract (English)
Electric guitar tone transfer (EGTT) and tone removal (EGTR) are two fundamental tasks in guitar tone modeling: EGTT replaces a recording's tone with that of a reference, while EGTR recovers the dry direct-input (DI) signal from a wet, processed recording. Despite their highly related nature, prior work has addressed them independently, and both works have yet to achieve satisfactory results. In this paper, we propose EG-VAE, a unified framework that jointly models EGTT and EGTR by disentangling frame-level content and global tone representations from wet recordings with a variational autoencoder. EGTT is achieved by recombining a source's content with a reference's tone, while EGTR is attained by a novel tone masking objective that enforces content-tone disentanglement during training and realizes the removal procedure at inference. To improve transfer to tones unseen in training, a second training stage shapes a smooth tone space through variational sampling and audio-effects augmentation. Experimental results from both objective and subjective evaluations demonstrate that EG-VAE outperforms task-specific baselines on transfer and removal. Demos are available at https://guitar-tone-demo.vercel.app/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。