arXiv:2409.18006cs.CL2024-09被引 12

多语言长文本模型在低资源语种上表现大幅下降,尤其面对多个目标句时。

Evaluating Multilingual Long-Context Models for Retrieval and Reasoning

  • 构建mLongRR数据集,跨五种语言评估多目标长文本检索与推理能力
  • 英语模型最高准确率96%,索马里语降至36%,三目标时英语仅40%
  • 揭示低资源语言和复杂上下文是当前模型的关键瓶颈,适合多语言研究者参考

近期大语言模型(LLMs)在处理长上下文方面表现出色,部分模型在合成检索任务中几乎实现完美记忆。然而,这些评估主要集中在英文,且上下文中仅包含一个目标句子。本文研究了多语言环境下模型性能的泛化能力,特别是多个目标句子的情况。我们构建了一个新数据集mLongRR,涵盖英语、越南语、印尼语、斯瓦希里语和索马里语五种语言,均使用拉丁字母但属于不同语系和资源水平。分析显示,语言间存在显著性能差距:最佳模型如Gemini-1.5和GPT-4o,在单目标下英语准确率约96%,索马里语仅36%;当目标句增至三个时,英语准确率降至40%,索马里语为0%。结果表明,长上下文处理、目标数量增加或低资源语言都会显著影响模型表现。

原文摘要 · Abstract (English)

Recent large language models (LLMs) demonstrate impressive capabilities in handling long contexts, some exhibiting near-perfect recall on synthetic retrieval tasks. However, these evaluations have mainly focused on English text and involved a single target sentence within lengthy contexts. Our work investigates how LLM performance generalizes to multilingual settings with multiple hidden target sentences. We create a new dataset -- mLongRR -- to comprehensively evaluate several multilingual long-context LLMs on retrieval and reasoning tasks across five languages: English, Vietnamese, Indonesian, Swahili, and Somali. These languages share the Latin script but belong to distinct language families and resource levels. Our analysis reveals a significant performance gap between languages. The best-performing models such as Gemini-1.5 and GPT-4o, achieve around 96% accuracy in English to around 36% in Somali with a single target sentence. However, this accuracy drops to 40% in English and 0% in Somali when dealing with three target sentences. Our findings highlight the challenges long-context LLMs face when processing longer contexts, an increase in the number of target sentences, or languages of lower resource levels.

多语言长文本检索低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。