arXiv:2410.18436cs.CL2024-10EMNLP被引 6

代码混用能激活大模型中的语言特定知识,提升低资源语言任务表现。

Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching

  • 构建英文-韩文混用数据集EnKoQA,研究代码混用对模型知识激活的影响。
  • 相比纯英文,混用文本更能激活模型中与韩语相关的领域知识。
  • 适合关注多语言模型、低资源语言处理的研究者和开发者。

近期大型语言模型(LLMs)展现出多语言能力,但其仍以英语为主,因训练语料中英语占主导。低资源语言的数据匮乏仍是关键挑战。代码混用(CS)是多语言使用者在对话中交替使用语言的现象,可传递微妙的文化与语言细节,这些在翻译中容易丢失,并能激发人类交流中的语言特定知识。为此,我们研究代码混用是否能激活或识别并利用大模型中用于解决低资源语言任务的知识。为支持研究,我们首先提出一个合成的英文-韩文代码混用问答数据集EnKoQA。通过将激活过程分解为知识识别与知识利用两个阶段,对多种多语言大模型进行了综合分析。结果表明,相较于纯英文文本,代码混用能更真实地激活模型内部的知识,尤其在语言特定领域表现显著,显示出代码混用在低资源语言任务中的潜在价值。

原文摘要 · Abstract (English)

Recent large language models (LLMs) demonstrate multilingual abilities, yet they are English-centric due to dominance of English in training corpora. The limited resource for low-resource languages remains a crucial challenge. Code-switching (CS), a phenomenon where multilingual speakers alternate between languages in a discourse, can convey subtle cultural and linguistic nuances that can be otherwise lost in translation and elicits language-specific knowledge in human communications. In light of this, we investigate whether code-switching can activate, or identify and leverage knowledge for reasoning when LLMs solve low-resource language tasks. To facilitate the research, we first present EnKoQA, a synthetic English-Korean CS question-answering dataset. We provide comprehensive analysis on a variety of multilingual LLMs by subdividing activation process into knowledge identification and knowledge leveraging. Our results demonstrate that compared to English text, CS can faithfully activate knowledge inside LLMs especially on language-specific domains, suggesting the potential of code-switching on low-resource language tasks.

多语言模型代码混用低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。