arXiv:2502.04269cs.CLcs.AI2025-02被引 2

探究多语言模型如何处理不同语言,揭示其对低资源语言的短板

How does a Multilingual LM Handle Multiple Languages?

  • 用词向量相似度分析多语言语义一致性
  • 发现高资源语言表现好,低资源语言识别能力弱
  • 适合关注多语言模型公平性与改进的研究者

多语言语言模型(MLMs)在自然语言处理快速发展的推动下取得显著进展,如基于多样化多语言数据集训练的BLOOM 1.7B和Qwen2,旨在弥合语言鸿沟。然而,它们在捕捉语言知识,尤其是低资源语言方面的能力仍不明确。本研究从多语言理解、语义表征和跨语言知识迁移三方面评估其性能。通过词向量余弦相似度分析语义一致性,利用命名实体识别和句子相似性任务考察语言结构,再以情感分析和文本分类任务测试从高资源语言向低资源语言的知识迁移能力。研究结合语言探针、性能指标与可视化方法,揭示了模型在高资源语言上表现良好,但在低资源语言中存在明显不足。结果有助于优化多语言NLP模型,提升对各类语言的支持能力,推动语言技术的包容性发展。

原文摘要 · Abstract (English)

Multilingual language models have significantly advanced due to rapid progress in natural language processing. Models like BLOOM 1.7B, trained on diverse multilingual datasets, aim to bridge linguistic gaps. However, their effectiveness in capturing linguistic knowledge, particularly for low-resource languages, remains an open question. This study critically examines MLMs capabilities in multilingual understanding, semantic representation, and cross-lingual knowledge transfer. While these models perform well for high-resource languages, they struggle with less-represented ones. Additionally, traditional evaluation methods often overlook their internal syntactic and semantic encoding. This research addresses key limitations through three objectives. First, it assesses semantic similarity by analyzing multilingual word embeddings for consistency using cosine similarity. Second, it examines BLOOM-1.7B and Qwen2 through Named Entity Recognition and sentence similarity tasks to understand their linguistic structures. Third, it explores cross-lingual knowledge transfer by evaluating generalization from high-resource to low-resource languages in sentiment analysis and text classification. By leveraging linguistic probing, performance metrics, and visualizations, this study provides insights into the strengths and limitations of MLMs. The findings aim to enhance multilingual NLP models, ensuring better support for both high- and low-resource languages, thereby promoting inclusivity in language technologies.

多语言模型低资源语言语义表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。