轻量级阿拉伯语语音模型,高效适配资源受限场景
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
- 用迭代自蒸馏训练阿拉伯语专用语音模型,保留语言特异性表征
- 在阿拉伯语语音识别、情感识别等任务上性能达顶尖水平
- 模型极小,适合移动端或低算力设备部署,适合中东地区研究者
大型预训练语音模型在下游任务中表现优异,但在资源受限环境中难以部署。本文提出 HArnESS,首个面向阿拉伯语的自监督语音模型系列,旨在捕捉阿拉伯语语音特征。通过迭代自蒸馏,我们训练了大型双语教师模型(HL),并将其知识压缩至轻量学生模型(HS、HST),保留阿拉伯语特异性表示。进一步采用低秩近似,将教师模型的离散监督信息压缩为浅层、薄型结构。我们在阿拉伯语语音识别(ASR)、说话人情绪识别(SER)和方言识别(DID)任务上评估,结果表明其性能优于 HuBERT 与 XLS-R。仅需少量微调,即达到最优或相当水平,是实际应用中的轻量高效替代方案。我们已公开蒸馏模型与研究成果,支持低资源环境下的负责任研究与部署。
原文摘要 · Abstract (English)
Large pre-trained speech models excel in downstream tasks but their deployment is impractical for resource-limited environments. In this paper, we introduce HArnESS, the first Arabic-centric self-supervised speech model family, designed to capture Arabic speech nuances. Using iterative self-distillation, we train large bilingual HArnESS (HL) SSL models and then distill knowledge into compressed student models (HS, HST), preserving Arabic-specific representations. We use low-rank approximation to further compact the teacher's discrete supervision into shallow, thin models. We evaluate HArnESS on Arabic ASR, Speaker Emotion Recognition (SER), and Dialect Identification (DID), demonstrating effectiveness against HuBERT and XLS-R. With minimal fine-tuning, HArnESS achieves SOTA or comparable performance, making it a lightweight yet powerful alternative for real-world use. We release our distilled models and findings to support responsible research and deployment in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。