arXiv:2606.30356cs.CLcs.LG2026-06

通过统一目标联合优化语音重建与视角增强的隐变量预测。

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

论文配图:OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL
图 1 · 摘自论文原文
  • 用多视角掩码隐变量预测+波形重建,统一训练目标。
  • 生成和说话人任务性能显著提升,识别与语义任务保持竞争力。
  • 适合语音生成、说话人识别等需要鲁棒表征的任务场景。

我们提出在线隐变量预测与不变视角及重建(OLIVE),一种自监督语音表示学习框架,联合优化分析与合成目标。OLIVE 在统一目标下结合了视角增强的掩码隐变量预测与波形重建。波形重建约束早期编码器特征保留信号级信息,而掩码隐变量预测则使后期上下文表示趋向不变性,以提升下游任务鲁棒性。实验表明,该框架支持多种任务。具体而言,OLIVE 在生成和说话人任务上表现更优,在识别与语义任务上保持竞争力,并提升了波形重建效果。

原文摘要 · Abstract (English)

We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives. OLIVE combines view-augmented masked latent prediction with waveform reconstruction under a unified objective. Reconstruction constrains early encoder features to retain signal-level information, while masked latent prediction shapes later contextual representations toward invariance for robust downstream performance. We show that these objectives enable representations that support a broad range of tasks. In particular, OLIVE improves results on generation and speaker tasks, maintains competitive performance on recognition and semantic tasks, and improves waveform reconstruction.

自监督学习语音表征波形重建说话人识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。