arXiv:2409.13897cs.CLcs.AI2024-09被引 5

提升低资源语言的LLM表现,让模型更包容多元文化。

LLM for Everyone: Representing the Underrepresented in Large Language Models

  • 用少量数据和计算量,优化低资源语言的模型性能
  • 在不损失通用任务能力的前提下,增强对小语种的泛化能力
  • 提出衡量跨语言文化价值观对齐的新方法,促进模型包容性

自然语言处理领域中,大语言模型在多种任务上表现出色,但在多语言场景下,尤其是低资源语言方面仍存在显著局限。本论文聚焦于低资源语言的研究空白,全面评估了大语言模型在这些语言中的表现,揭示了多语言与跨文化泛化能力的不足。为弥合这一差距,本文提出了数据与计算高效的解决方案,包括跨语言持续指令微调、基于检索的跨语言上下文学习,以及上下文查询对齐方法,有效提升了模型在低资源语言上的泛化能力,同时保持原有任务泛化性能不受影响。此外,提出一种新方法以度量不同语言环境下模型的文化价值观一致性,确保模型具备文化敏感性与包容性。这些贡献推动大语言模型在多语言与跨文化场景下的公平性与可及性,助力自然语言处理向更均衡、包容的方向发展。

原文摘要 · Abstract (English)

Natural language processing (NLP) has witnessed a profound impact of large language models (LLMs) that excel in a multitude of tasks. However, the limitation of LLMs in multilingual settings, particularly in underrepresented languages, remains a significant hurdle. This thesis aims to bridge the gap in NLP research and development by focusing on underrepresented languages. A comprehensive evaluation of LLMs is conducted to assess their capabilities in these languages, revealing the challenges of multilingual and multicultural generalization. Addressing the multilingual generalization gap, this thesis proposes data-and-compute-efficient methods to mitigate the disparity in LLM ability in underrepresented languages, allowing better generalization on underrepresented languages without the loss of task generalization ability. The proposed solutions cover cross-lingual continual instruction tuning, retrieval-based cross-lingual in-context learning, and in-context query alignment. Furthermore, a novel method to measure cultural values alignment between LLMs operating in different languages is proposed, ensuring cultural sensitivity and inclusivity. These contributions aim to enhance the multilingual and multicultural alignment of LLMs in underrepresented languages, ultimately advancing the NLP field toward greater equality and inclusiveness.

大模型多语言公平性文化对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。