arXiv:2605.13084cs.CLcs.AI2026-05

多语言语音词分类中,数据量比语言种类更重要。

Does language matter for spoken word classification? A multilingual generative meta-learning approach

论文配图:Does language matter for spoken word classification? A multilingual generative meta-learning approach
图 1 · 摘自论文原文
  • 用生成式元学习方法训练跨语言语音模型
  • 多语言模型表现最优,但各模型差距小
  • 训练时看到的独特语音数据时长决定性能

元学习在少样本单语语音词分类中表现优于监督学习,但在多语言场景下仍研究不足。本文采用生成式元持续学习算法进行语音词分类。该算法具备生成能力,适合实际应用,元学习特性有助于跨语言泛化。我们分别在英语、德语、法语和加泰罗尼亚语上训练单语模型,英语与德语上训练双语模型,四语言联合训练多语言模型。结果发现,尽管多语言模型性能最佳,但不同模型间差异出人意料地小。更重要的是,训练中接触的唯一语音数据时长比语言数量更能预测模型表现。

原文摘要 · Abstract (English)

Meta-learning has been shown to have better performance than supervised learning for few-shot monolingual spoken word classification. However, the meta-learning approach remains under-explored in multilingual spoken word classification. In this paper, we apply the Generative Meta-Continual Learning algorithm to spoken word classification. The generative nature of this algorithm makes it viable for use in application, and the meta-learning aspect promotes generalisation, which is crucial in a multilingual setting. We train monolingual models on English, German, French, and Catalan, a bilingual model on English and German, and a multilingual model on all four languages. We find that although the multilingual model performs best, the differences between model performance is unexpectedly low. We also find that the hours of unique data seen during training seems to be a stronger performance indicator than the number of languages included in the training data.

语音分类元学习多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。