将开源语音模型适配至新加坡多语言执法场景,性能超越更大模型。
Efficiently Adapting Spoken Language Models for the Singaporean Context

- 用LoRA微调+替代数据集防遗忘,多任务学习优化语音理解。
- 在5个任务上表现超7倍大的模型,语音问答能力损失不足2%。
- 适合需要多语言语音交互的政府/安全系统部署应用。
语音语言模型(SLMs)统一了语音感知与推理能力,但在敏感领域中的适配研究仍不充分,尤其当原始训练数据不可获取且需支持多语言语音查询时。本文将一个开源SLM适配至新加坡执法部门(Home Team)场景,在新加坡四种官方语言下完成五项语音任务。方法结合LoRA微调、用于防止灾难性遗忘的代理文本问答数据集,以及针对语音优化的CoBa重加权多任务目标。同时构建了包含504,853条样本的多语言问答数据集HTD-multilingual-QA(文本与语音形式)。最终模型HT-Moonstone(5B)在多数任务上表现匹配或优于7倍大小的SLM,语音识别中对口音和性别辨识最佳,语音问答能力仅损失不到2%。
原文摘要 · Abstract (English)
Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spoken-query interaction. We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore's four official languages, combining LoRA fine-tuning, a surrogate text-QA dataset that guards against catastrophic forgetting, and a multi-task objective that adapts the CoBa reweighting scheme to speech. We also build HTD-multilingual-QA, a 504,853 sample multilingual QA dataset in text and spoken form. The resulting HT-Moonstone (5B) matches or outperforms SLMs up to 7x its size on most tasks, attains the best accent and gender recognition among all models evaluated, and loses under 2\% of its original speech QA ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。