arXiv:2508.17229cs.SDcs.AI2025-08AAAI被引 7

用多指标偏好对齐提升语音修复生成质量,避免感知偏差。

Multi-Metric Preference Alignment for Generative Speech Restoration

  • 构建80K条多维度偏好数据对,融合感知、保真、内容与音色指标
  • 在三种生成模型上均显著提升客观与主观效果,验证方法普适性
  • 可作伪标签生成器,助力稀缺数据场景下的模型训练

近期生成模型显著推动了语音修复任务的发展,但其训练目标常与人类感知偏好不一致,导致质量不佳。尽管后训练偏好对齐在文本和图像生成中已证明有效,但在语音修复领域仍研究不足。本文探讨该方法的挑战,重点解决如何定义稳健的偏好信号及构建高质量数据以避免奖励欺骗。为此提出多指标偏好对齐策略,构建新数据集GenSR-Pref,包含80万组偏好对,每组中被选样本均被一组涵盖感知质量、信号保真度、内容一致性与音色保留的互补指标一致青睐。该原则化方法确保偏好信号全面性。使用该数据集进行直接偏好优化(DPO),在三种不同生成范式——自回归模型(AR)、掩码生成模型(MGM)与流匹配模型(FM)——上,在多个修复基准测试中均实现一致且显著的性能提升,涵盖客观与主观评估。消融实验表明,多指标策略相比单指标更有效缓解奖励欺骗。此外,实证表明对齐模型可作为强大‘数据标注者’,在数据稀缺场景(如歌声修复)下生成高质量伪标签,为传统判别模型提供监督信号。

原文摘要 · Abstract (English)

Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptimal quality. While post-training alignment has proven effective in other generative domains like text and image generation, its application to generative speech restoration remains largely under-explored. This work investigates the challenges of applying preference-based post-training to this task, focusing on how to define a robust preference signal and curate high-quality data to avoid reward hacking. To address these challenges, we propose a multi-metric preference alignment strategy. We construct a new dataset, GenSR-Pref, comprising 80K preference pairs, where each chosen sample is unanimously favored by a complementary suite of metrics covering perceptual quality, signal fidelity, content consistency, and timbre preservation. This principled approach ensures a holistic preference signal. Applying Direct Preference Optimization (DPO) with our dataset, we observe consistent and significant performance gains across three diverse generative paradigms: autoregressive models (AR), masked generative models (MGM), and flow-matching models (FM) on various restoration benchmarks, in both objective and subjective evaluations. Ablation studies confirm the superiority of our multi-metric strategy over single-metric approaches in mitigating reward hacking. Furthermore, we demonstrate that our aligned models can serve as powerful ''data annotators'', generating high-quality pseudo-labels to serve as a supervision signal for traditional discriminative models in data-scarce scenarios like singing voice restoration. Demo Page:https://gensr-pref.github.io

语音修复生成模型偏好对齐数据标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。