用多模态反馈提升歌声合成评估的可解释性
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
- 用音频-语言模型生成文本和音频批评,覆盖旋律、内容等维度
- 在混合数据集上微调,提升反馈多样性与语言丰富性
- 生成结果音乐准确且可解释,适合指导生成模型优化
歌声合成(SVS)已取得显著进展,能够生成音高准确、风格一致的演唱。随着能力提升,可靠评估与优化的需求日益迫切。然而,现有方法如奖励系统常依赖单一数值评分,难以捕捉乐句、表现力等多维特性,且需昂贵标注,影响可解释性与泛化能力。为此,我们提出一种生成式反馈框架(即奖励模型),提供多维度的语言与音频反馈用于SVS评估。该方法利用音频-语言模型生成涵盖旋律、内容与听觉质量等方面的文本与音频批评。模型在混合数据集上微调,该数据集结合了人类音乐反应与来自多模态大语言模型(MLLMs)的合成批评,增强了反馈的多样性与语言丰富性。定量实验验证了所提数据集与训练策略的有效性,表明该框架能生成音乐准确且可解释的评估结果,适用于指导生成模型改进。代码已公开于 https://github.com/opendilab/VocalCritic。
原文摘要 · Abstract (English)
Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for reliable evaluation and optimization becomes increasingly critical. However, current methods like reward systems often rely on single numerical scores, struggle to capture various dimensions such as phrasing or expressiveness, and require costly annotations, limiting interpretability and generalization. To address these issues, we propose a generative feedback (i.e., reward model) framework that provides multi-dimensional language and audio feedback for SVS assessment. Our approach leverages an audio-language model to generate text and audio critiques-covering aspects such as melody, content, and auditory quality. The model is fine-tuned on a hybrid dataset combining human music reactions and synthetic critiques from a MLLMs, enhancing diversity and linguistic richness. Quantitative experiments validate the effectiveness of the proposed dataset and training strategy, demonstrating that the framework produces musically accurate and interpretable evaluations suitable for guiding generative model improvement. The code is at [https://github.com/opendilab/VocalCritic](https://github.com/opendilab/VocalCritic)
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。