让语音模型自动感知语言和说话人,提升跨任务泛化能力
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
- 通过动态调制早期嵌入,让自监督模型感知语言与说话人条件
- 在未见任务上显著降低错误率,如说话人识别错误降10%
- 参数少、不易过拟合,适合资源不足场景
我们提出条件感知的自监督学习表征(CA-SSLR),一种适用于多种语音处理任务的通用条件模型。与传统微调方法不同,CA-SSLR在早期层融合语言与说话人嵌入,使SSL模型具备当前语言和说话人上下文意识。该方法降低对输入音频特征的依赖,同时保持基础SSL表征完整性。通过线性调制实现内部表示的动态调整,可在不显著改变原模型行为的前提下实现细粒度适应。实验表明,CA-SSLR减少可训练参数,缓解过拟合,在低资源及未见任务中表现优异:在说话人识别任务中相对误差降低10%,在ML-SUPERB基准上自动语音识别词错误率下降37%,在VoxCeleb-1上说话人验证等错误率降低27%。
原文摘要 · Abstract (English)
We introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to standard fine-tuning methods that optimize for downstream models, CA-SSLR integrates language and speaker embeddings from earlier layers, making the SSL model aware of the current language and speaker context. This approach reduces the reliance on input audio features while preserving the integrity of the base SSLR. CA-SSLR improves the model's capabilities and demonstrates its generality on unseen tasks with minimal task-specific tuning. Our method employs linear modulation to dynamically adjust internal representations, enabling fine-grained adaptability without significantly altering the original model behavior. Experiments show that CA-SSLR reduces the number of trainable parameters, mitigates overfitting, and excels in under-resourced and unseen tasks. Specifically, CA-SSLR achieves a 10% relative reduction in LID errors, a 37% improvement in ASR CER on the ML-SUPERB benchmark, and a 27% decrease in SV EER on VoxCeleb-1, demonstrating its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。