arXiv:2411.15577cs.CL2024-11被引 4

用语法描述增强模型,提升低资源语言的零样本分类能力

From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars

  • 基于检索增强生成,从语法描述中提取关键信息
  • 覆盖248种语言、142个语系,评估模型在复杂语法中的理解力
  • 为低资源语言的NLP研究提供首个系统性评测基准

语言建模的最新进展显著提升了零样本能力,包括上下文学习、指令遵循及极低资源语言的机器翻译(Tanzer等,2024)。然而,许多书写资源有限的语言主要依赖形式化的语法与词汇描述。本文提出一套新基准,评估模型从复杂语言语法中提取和分类信息的能力。我们采用基于检索增强生成(RAG)的方法,利用这些语法描述支持下游任务如机器翻译。该基准涵盖来自142个语系的248种语言,聚焦WALS与Grambank中的类型学特征。这是首个全面评估语言模型在上下文中准确解读与提取语言特征的能力的资源,对推动低资源语言的NLP发展具有重要意义。代码与数据已公开于\url{https://github.com/al-the-eigenvalue/RAG-on-grammars}。

原文摘要 · Abstract (English)

Recent advances in language modeling have demonstrated significant improvements in zero-shot capabilities, including in-context learning, instruction following, and machine translation for extremely under-resourced languages (Tanzer et al., 2024). However, many languages with limited written resources rely primarily on formal descriptions of grammar and vocabulary. In this paper, we introduce a set of benchmarks to evaluate how well models can extract and classify information from the complex descriptions found in linguistic grammars. We present a Retrieval-Augmented Generation (RAG)-based approach that leverages these descriptions for downstream tasks such as machine translation. Our benchmarks encompass linguistic descriptions for 248 languages across 142 language families, focusing on typological features from WALS and Grambank. This set of benchmarks offers the first comprehensive evaluation of language models' in-context ability to accurately interpret and extract linguistic features, providing a critical resource for scaling NLP to low-resource languages. The code and data are publicly available at \url{https://github.com/al-the-eigenvalue/RAG-on-grammars}.

语言模型低资源语言语法解析检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。