跨语言语音性别识别模型,支持英语与东南亚多语种
MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages
- 基于大模型微调,用低秩适配提升效率
- 在多语种场景下性能超越现有顶尖模型
- 适合需要跨语言语音分析的科研与应用
我们提出MERaLiON-GR,一种针对英语和东南亚(SEA)语言的语音性别识别系统,实现男女二分类。该模型在大规模语音语料预训练的Conformer架构基础上,采用低秩适配(LoRA)进行参数高效微调,并附加多尺度ECAPA-TDNN下游网络,结合注意力池化与轻量线性分类器。在新加坡及东南亚多语言(英语、中文、马来语、泰米尔语、泰语、越南语、印尼语、高棉语)上的大量评估显示,该模型在全句与分段两种评估模式下均显著优于当前最优的Vox-Profile模型和大型音频大模型。结果表明,专用语音模型对实现精准的副语言理解与强跨语言泛化具有重要价值。
原文摘要 · Abstract (English)
We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes MERaLiON-SpeechEncoder-2, a large conformer based transformer pre-trained on a broad speech corpus, and applies parameter efficient fine-tuning via Low-Rank Adaptation (LoRA) to adapt the encoder to the gender recognition task, and appends a multi-scale ECAPA-TDNN down stream network with attention pooling and a lightweight linear classifier. Extensive evaluations across multilingual Singaporean and Southeast Asian languages (English, Chinese, Malay, Tamil, Thai, Vietnamese, Indonesian, and Khmer) show that MERaLiON-GR consistently surpasses the state-of-the-art gender recognition model Vox-Profile and a large Audio-LLM, in both full-utterance and segment level evaluation modes. The results underscore the value of dedicated speech models in achieving accurate paralinguistic understanding and strong cross-lingual generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。