用自然语言控制,一键修复音乐各类质量问题。
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
- 基于文本指令的统一模型,可精准修复多种音频缺陷。
- 在19种退化函数下,客观指标显著提升音质。
- 适合非专业用户快速获得高质量音乐成品。
音乐录音常因非专业环境录制而出现混响过强、失真、削波、音色不平衡和立体声范围狭窄等问题。传统解决方式依赖多个专用工具和手动调整。本文提出SonicMaster,首个统一的生成式音乐修复与母带处理模型,支持自然语言指令控制,可针对性优化或自动修复。为训练该模型,构建了包含19种退化函数(分属均衡、动态、混响、幅度、立体声五类)的SonicMaster数据集。采用流匹配生成训练范式,学习从退化输入到高质量输出的映射关系,由文本提示引导。客观评估显示,SonicMaster在所有退化类型上均显著提升音质;主观听感测试表明,听众更偏好其输出结果。
原文摘要 · Abstract (English)
Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment or expertise. These problems are typically corrected using separate specialized tools and manual adjustments. In this paper, we introduce SonicMaster, the first unified generative model for music restoration and mastering that addresses a broad spectrum of audio artifacts with text-based control. SonicMaster is conditioned on natural language instructions to apply targeted enhancements, or can operate in an automatic mode for general restoration. To train this model, we construct the SonicMaster dataset, a large dataset of paired degraded and high-quality tracks by simulating common degradation types with nineteen degradation functions belonging to five enhancements groups: equalization, dynamics, reverb, amplitude, and stereo. Our approach leverages a flow-matching generative training paradigm to learn an audio transformation that maps degraded inputs to their cleaned, mastered versions guided by text prompts. Objective audio quality metrics demonstrate that SonicMaster significantly improves sound quality across all artifact categories. Furthermore, subjective listening tests confirm that listeners prefer SonicMaster's enhanced outputs over other baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。