用灵活对齐学习语音表示,提升不同语速下的鲁棒性。
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
- 结合孪生网络与CTC损失,实现内容一致的非严格帧对齐
- 在不同语速下表现更优,下游任务性能显著提升
- 适合需要抗语速变化的语音识别与表示学习场景
自监督语音表示学习通过孪生网络取得了显著进展,该方法利用同一输入的不同视图进行训练。然而,现有方法通常要求这些视图间存在严格的帧级对齐,忽略了不同说话风格下的语言上下文不变性。本文提出SiamCTC框架,将孪生网络与连接时序分类(CTC)相结合,无需严格帧对齐即可学习语音表示。通过CTC损失建立相同内容在不同时间尺度上的柔性单调对齐,有效应对语速扰动及其他时间增强。该设计放宽了帧级约束,同时保持时间连贯性,增强了下游任务对语速变化的鲁棒性。实验表明,SiamCTC能生成更具适应性的语音表示,尤其在多样语速下表现优异。
原文摘要 · Abstract (English)
Self-supervised speech representation learning has made significant progress through Siamese networks, which leverage different views of the same input. However, existing methods often require frame-wise alignment between these views, overlooking the broader linguistic context invariance across different speaking styles. We introduce SiamCTC, a framework that integrates Siamese networks with Connectionist Temporal Classification (CTC) to learn speech representations without strict frame-level correspondence. By employing CTC loss to establish flexible, monotonic alignments between differing temporal realizations of the same content, SiamCTC accommodates speed perturbations and other temporal augmentations. This design relaxes frame-wise constraints while preserving temporal coherence and enhancing robustness to speaking-rate variations in downstream tasks. Our experiments demonstrate that SiamCTC leads to more adaptable speech representations, particularly at diverse speaking rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。