多语言大模型或可为欧洲小语种提供技术出路,助力数字语言平等
Are Multilingual Language Models an Off-ramp for Under-resourced Languages? Will we arrive at Digital Language Equality in Europe in 2030?
- 利用多语言数据训练的LLM,可提升低资源语言表现
- 当前技术已支持部分欧洲小语种的自然语言处理任务
- 适合关注语言公平与多语种AI落地的研究者
大型语言模型(LLMs)在几乎所有自然语言处理(NLP)任务和语言技术(LT)应用中表现出前所未有的能力,但其训练依赖充足预训练数据,导致多数低资源语言被排除在外。然而,已有间接和实证证据表明,基于多语言数据(含低资源语言)训练的多语言LLM,对部分低资源语言也展现出较强能力。这一方法可能成为无法发展‘原生’LLM的低资源语言的技术突破路径。本文聚焦欧洲语言,评估现有技术支撑状况,综述相关研究,并指出实现该路径需解决的关键开放问题。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate unprecedented capabilities and define the state of the art for almost all natural language processing (NLP) tasks and also for essentially all Language Technology (LT) applications. LLMs can only be trained for languages for which a sufficient amount of pre-training data is available, effectively excluding many languages that are typically characterised as under-resourced. However, there is both circumstantial and empirical evidence that multilingual LLMs, which have been trained using data sets that cover multiple languages (including under-resourced ones), do exhibit strong capabilities for some of these under-resourced languages. Eventually, this approach may have the potential to be a technological off-ramp for those under-resourced languages for which "native" LLMs, and LLM-based technologies, cannot be developed due to a lack of training data. This paper, which concentrates on European languages, examines this idea, analyses the current situation in terms of technology support and summarises related work. The article concludes by focusing on the key open questions that need to be answered for the approach to be put into practice in a systematic way.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。