针对真实场景下语音编辑的音质不连贯问题,提出测试时自适应优化方法。
Instance-Specific Test-Time Training for Speech Editing in the Wild
- 通过真实语音特征直接监督未编辑区域,间接约束编辑区域。
- 在野外数据集上主观评价得分提升15%,客观指标显著优于现有方法。
- 适合需要高保真语音编辑的实时应用,如语音助手、播客制作。
语音编辑系统旨在自然修改语音内容的同时保持声学一致性和说话人身份。然而,以往研究在面对未见过且多样的声学条件时往往表现不佳,导致真实场景下编辑性能下降。为此,我们提出一种面向真实场景的实例特定测试时训练方法。该方法利用未编辑区域的真实声学特征进行直接监督,并通过基于持续时间约束和音素预测的辅助损失,在编辑区域实现间接监督。这一策略有效缓解了语音编辑中的带宽不连续问题,确保未编辑与编辑区域之间的声学过渡平滑。同时,通过测试时训练中调整掩码长度,可精确控制语速以适配目标时长。在野外基准数据集上的实验表明,本方法在客观和主观评估中均优于现有语音编辑系统。
原文摘要 · Abstract (English)
Speech editing systems aim to naturally modify speech content while preserving acoustic consistency and speaker identity. However, previous studies often struggle to adapt to unseen and diverse acoustic conditions, resulting in degraded editing performance in real-world scenarios. To address this, we propose an instance-specific test-time training method for speech editing in the wild. Our approach employs direct supervision from ground-truth acoustic features in unedited regions and indirect supervision in edited regions via auxiliary losses based on duration constraints and phoneme prediction. This strategy mitigates the bandwidth discontinuity problem in speech editing, ensuring smooth acoustic transitions between unedited and edited regions. Additionally, it enables precise control over speech rate by adapting the model to target durations via mask length adjustment during test-time training. Experiments on in-the-wild benchmark datasets demonstrate that our method outperforms existing speech editing systems in both objective and subjective evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。