评测主流大模型在粤语、日语和土耳其语上的表现,揭示其跨语言能力短板。
Evaluating Modern Large Language Models on Low-Resource and Morphologically Rich Languages:A Cross-Lingual Benchmark Across Cantonese, Japanese, and Turkish
- 构建涵盖三语的跨语言评测基准,覆盖问答、摘要等四类任务
- 大模型在文化语境理解上普遍不足,小模型差距更大
- 开源模型在方言和形态复杂语言中表现明显落后
大型语言模型(LLMs)在英语等高资源语言上表现优异,但在低资源且形态丰富的语言中仍缺乏深入研究。本文对七款前沿LLM——GPT-4o、GPT-4、Claude 3.5 Sonnet、LLaMA 3.1、Mistral Large 2、LLaMA-2 Chat 13B和Mistral 7B Instruct——在新提出的跨语言基准上进行评估,涵盖粤语、日语和土耳其语。该基准包含开放域问答、文档摘要、英译多语种及文化情境对话四类任务。采用人工评分(流畅性、事实准确性和文化适宜性)与自动指标(如BLEU、ROUGE)结合的方式评估性能。结果显示,最大型专有模型(如GPT-4o、GPT-4、Claude 3.5)整体领先,但在文化语境理解和形态泛化方面仍存在显著差距。其中,GPT-4o在跨语言任务中表现稳健,Claude 3.5 Sonnet在知识推理任务中达到竞争性准确率。然而,所有模型均难以应对土耳其语的黏着构词和粤语口语表达等独特挑战。较小的开源模型(如LLaMA-2 13B、Mistral 7B)在流畅度和准确性上显著落后,凸显资源不平等。我们提供详细定量结果与定性错误分析,并讨论如何开发更具文化敏感性和语言泛化能力的模型。相关基准与数据已公开以促进可复现性与后续研究。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved impressive results in high-resource languages like English, yet their effectiveness in low-resource and morphologically rich languages remains underexplored. In this paper, we present a comprehensive evaluation of seven cutting-edge LLMs -- including GPT-4o, GPT-4, Claude~3.5~Sonnet, LLaMA~3.1, Mistral~Large~2, LLaMA-2~Chat~13B, and Mistral~7B~Instruct -- on a new cross-lingual benchmark covering \textbf{Cantonese, Japanese, and Turkish}. Our benchmark spans four diverse tasks: open-domain question answering, document summarization, English-to-X translation, and culturally grounded dialogue. We combine \textbf{human evaluations} (rating fluency, factual accuracy, and cultural appropriateness) with automated metrics (e.g., BLEU, ROUGE) to assess model performance. Our results reveal that while the largest proprietary models (GPT-4o, GPT-4, Claude~3.5) generally lead across languages and tasks, significant gaps persist in culturally nuanced understanding and morphological generalization. Notably, GPT-4o demonstrates robust multilingual performance even on cross-lingual tasks, and Claude~3.5~Sonnet achieves competitive accuracy on knowledge and reasoning benchmarks. However, all models struggle to some extent with the unique linguistic challenges of each language, such as Turkish agglutinative morphology and Cantonese colloquialisms. Smaller open-source models (LLaMA-2~13B, Mistral~7B) lag substantially in fluency and accuracy, highlighting the resource disparity. We provide detailed quantitative results, qualitative error analysis, and discuss implications for developing more culturally aware and linguistically generalizable LLMs. Our benchmark and evaluation data are released to foster reproducibility and further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。