统一音频与语音表征学习,提升跨任务性能
ULTRAS -- Unified Learning of Transformer Representations for Audio and Speech Signals
- 在长片段上进行掩码与预测建模,融合时频特征
- 在多个语音和音频任务中优于现有基线模型
- 适合需要统一表征的多模态音频应用
自监督学习(SSL)通过时域预测目标在语音处理中取得显著进展,而音频表征学习框架则基于时频谱图。针对不同范式间迁移困难的问题,本文提出统一音频与语音表征学习框架ULTRAS。该模型基于Transformer架构,对对数梅尔谱图的频谱块进行编码,并在长片段上进行掩码与预测建模。通过联合损失函数对频谱和时间目标进行预测,迫使表示同时捕捉时序与频率特性。在多种语音与音频任务上的实验表明,ULTRAS框架性能优于现有基准模型。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has driven impressive advances in speech processing by adopting time-domain prediction objectives, while audio representation learning frameworks operate on time-frequency spectrograms. Models optimized for one paradigm struggle to transfer to the other, highlighting the need for a joint framework. We propose Unified Learning of Transformer Representations for Audio and Speech (ULTRAS), where the masking and predictive modeling is performed over long patches of the data. The model, based on the transformer architecture, encodes spectral-patches of log-mel spectrogram features. The predictive modeling of masked segments is performed on spectral and temporal targets using a combined loss-function, forcing the representations to encode time and frequency traits. Experiments are performed on a variety of speech and audio tasks, where we illustrate that the ULTRAS framework achieves improved performance over other established baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。