arXiv:2410.13153cs.CL2024-10被引 22

对比大模型在英语与南亚低资源语言的表现差异

Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings

  • 用零样本提示测试中英文及南亚语言下的模型表现
  • GPT-4 在所有语言和提示设置中均优于其他模型
  • 英语提示效果显著好于低资源语言,凸显多语支持不足

大型语言模型(LLMs)基于海量数据训练,已广泛应用于多种任务。尽管性能优异,但多数模型主要在英语环境下开发与评估。近年来虽出现少数多语言模型,但其在低资源语言(尤其是南亚地区主流语言如孟加拉语、印地语、乌尔都语)中的表现仍缺乏研究。本研究评估了 GPT-4、Llama 2 与 Gemini 在英语及其他南亚低资源语言上的表现,采用零样本提示与五种不同提示策略,系统考察跨语言翻译提示的有效性。结果表明:在所有提示设置与语言中,GPT-4 均优于 Llama 2 与 Gemini;三者在英语提示下的表现均显著高于其他语言。研究揭示了当前大模型在多语言场景下的局限性,强调需提升模型泛化能力与语言资源以推动通用自然语言处理发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are trained on massive amounts of data, enabling their application across diverse domains and tasks. Despite their remarkable performance, most LLMs are developed and evaluated primarily in English. Recently, a few multi-lingual LLMs have emerged, but their performance in low-resource languages, especially the most spoken languages in South Asia, is less explored. To address this gap, in this study, we evaluate LLMs such as GPT-4, Llama 2, and Gemini to analyze their effectiveness in English compared to other low-resource languages from South Asia (e.g., Bangla, Hindi, and Urdu). Specifically, we utilized zero-shot prompting and five different prompt settings to extensively investigate the effectiveness of the LLMs in cross-lingual translated prompts. The findings of the study suggest that GPT-4 outperformed Llama 2 and Gemini in all five prompt settings and across all languages. Moreover, all three LLMs performed better for English language prompts than other low-resource language prompts. This study extensively investigates LLMs in low-resource language contexts to highlight the improvements required in LLMs and language-specific resources to develop more generally purposed NLP applications.

大模型评估多语言低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。