arXiv:2606.19629cs.SDcs.AI2026-06

通过强制语音编辑模型自洽,提升对标注噪声的鲁棒性。

RIVET: Robust Idempotent Voice Attribute Editing

论文配图:RIVET: Robust Idempotent Voice Attribute Editing
图 1 · 摘自论文原文
  • 引入自洽性约束,使重复编辑结果不变
  • 在含噪数据上编辑成功率提升,说话人身份保留更好
  • 适合语音属性编辑与标注质量差的场景

语音属性编辑模型可在保持说话人身份的前提下修改年龄、性别等特征。然而,在大规模语音数据集中,属性标注常存在噪声或不一致,导致条件生成模型产生不稳定编辑结果。本文表明,自洽性(idempotency)可有效提升对标签噪声的鲁棒性。自洽算子指重复应用不改变输出,即 f(f(x)) = f(x)。强制该性质相当于隐式正则化,降低对误标样本的敏感性。我们提出 RIVET 训练框架,引入自洽性目标以增强鲁棒性。在受控噪声和 GLOBE 数据集(自然含噪标注)上评估,RIVET 在编辑成功率和说话人身份保留方面均优于标准训练,证明自洽性能显著提升语音编辑模型的鲁棒性。

原文摘要 · Abstract (English)

Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datasets, however, attribute annotations are often noisy or inconsistent, which can cause conditional generative models to produce unstable edits. In this work, we show that idempotency provides an effective mechanism for improving robustness to noisy labels. An idempotent operator is one for which repeated application does not change the result, i.e., f(f(x)) = f(x). Enforcing this property acts as an implicit regularizer that reduces sensitivity to mislabeled examples. We introduce RIVET, a training framework that incorporates an idempotency objective to improve robustness to label noise. We evaluate RIVET under controlled label noise and on the GLOBE dataset with naturally noisy annotations. RIVET improves editing success and better preserves speaker identity than standard training, showing that idempotency improves robustness in voice editing models.

语音编辑鲁棒性自洽性标注噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。