轻量级阿拉伯语语音模型,兼顾准确与效率。
HARNESS: Lightweight Distilled Arabic Speech Foundation Models

- 从头训练并迭代自蒸馏,保留阿拉伯语语音特征。
- 在ASR、方言识别等任务上优于HuBERT和XLS-R。
- 适合资源受限场景,适合做阿拉伯语语音应用。
大型自监督语音(SSL)模型在下游任务中表现优异,但其规模限制了在资源受限环境中的部署。我们提出HArnESS,一个从头开始训练的阿拉伯语中心自监督语音模型系列,采用迭代自蒸馏方法,并配备轻量级学生模型,在自动语音识别(ASR)、方言识别(DID)和语音情感识别(SER)任务上实现强准确率-效率平衡。方法从一个双语阿拉伯语-英语教师模型出发,逐步将知识蒸馏至压缩的学生模型,同时保留与阿拉伯语相关的声学和副语言表征。我们进一步研究基于PCA的教师监督信号压缩,以更好地匹配浅层薄层学生模型的容量。相较于HuBERT和XLS-R,HArnESS在阿拉伯语下游任务中持续提升性能,而压缩模型在大幅结构缩减下仍保持竞争力。这些结果使HArnESS成为现实世界语音应用中实用且可访问的阿拉伯语中心基础模型。
原文摘要 · Abstract (English)
Large self-supervised speech (SSL) models achieve strong downstream performance, but their size limits deployment in resource-constrained settings. We present HArnESS, an Arabic-centric self-supervised speech model family trained from scratch with iterative self-distillation, together with lightweight student variants that offer strong accuracy-efficiency trade-offs on Automatic Speech Recognition (ASR), Dialect Identification (DID), and Speech Emotion Recognition (SER). Our approach begins with a large bilingual Arabic-English teacher and progressively distills its knowledge into compressed student models while preserving Arabic-relevant acoustic and paralinguistic representations. We further study PCA-based compression of the teacher supervision signal to better match the capacity of shallow and thin students. Compared with HuBERT and XLS-R, HArnESS consistently improves performance on Arabic downstream tasks, while the compressed models remain competitive under substantial structural reduction. These results position HArnESS as a practical and accessible Arabic-centric SSL foundation for real-world speech applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。