研究大模型在古语言上的零样本泛化能力,发现模型规模是关键。
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
- 用检索增强生成提升梵语问答性能
- 小模型在古语言任务中表现明显下降
- 模型越大,跨语言泛化能力越强
大型语言模型在多种任务和语言上展现出卓越的泛化能力。本研究聚焦梵语、古希腊语和拉丁语这三种古典语言,探究影响跨语言零样本泛化的主要因素。首先,我们开展命名实体识别和向英语的机器翻译任务。尽管大模型在域外数据上表现优于或等同于微调基线,但小模型在特定或抽象实体类型上仍表现不佳。其次,我们以梵语为例构建了一个事实型问答(QA)数据集,发现引入检索增强生成可显著提升性能;相比之下,小模型在该类任务中出现明显性能下降。结果表明,模型规模是影响跨语言泛化的重要因素。即使未针对古典语言进行指令微调,GPT-4o 和 Llama-3.1 等模型仍能有效处理这些语言,为古典研究提供了新工具。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable generalization capabilities across diverse tasks and languages. In this study, we focus on natural language understanding in three classical languages -- Sanskrit, Ancient Greek and Latin -- to investigate the factors affecting cross-lingual zero-shot generalization. First, we explore named entity recognition and machine translation into English. While LLMs perform equal to or better than fine-tuned baselines on out-of-domain data, smaller models often struggle, especially with niche or abstract entity types. In addition, we concentrate on Sanskrit by presenting a factoid question-answering (QA) dataset and show that incorporating context via retrieval-augmented generation approach significantly boosts performance. In contrast, we observe pronounced performance drops for smaller LLMs across these QA tasks. These results suggest model scale as an important factor influencing cross-lingual generalization. Assuming that models used such as GPT-4o and Llama-3.1 are not instruction fine-tuned on classical languages, our findings provide insights into how LLMs may generalize on these languages and their consequent utility in classical studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。