为希腊语问答构建新数据集,评估中英文大模型表现差异。
Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark
- 构建社交媒体来源的希腊语问答数据集DemosQA,贴近本土文化。
- 在6个希腊语数据集上评估11个模型,发现单语模型更适配本地任务。
- 开源代码与数据,支持多语言评测框架复现。
近年来自然语言处理和深度学习的发展推动了大型语言模型(LLMs)的兴起,显著提升了各类任务的表现,包括问答(QA)。然而,现有研究主要集中在高资源语言(如英语),近期才关注多语言模型。这些模型往往偏向少数主流语言,或依赖从高资源语言到低资源语言的迁移学习,可能导致对社会、文化及历史背景的误读。为此,虽有针对低资源语言的单语模型被开发,但其在特定语言任务上的效果仍缺乏与多语言模型的系统对比。本研究以希腊语问答为切入点,贡献:(i) DemosQA,一个基于社交媒体用户提问与社区审核回答构建的新数据集,更真实反映希腊社会文化语境;(ii) 一种内存高效的LLM评估框架,可适配多种语言和数据集;(iii) 在6个人工筛选的希腊语问答数据集上,使用3种提示策略,对11个单语与多语言模型进行广泛评估。所有代码与数据均已公开,以促进研究可复现性。
原文摘要 · Abstract (English)
Recent advancements in Natural Language Processing and Deep Learning have enabled the development of Large Language Models (LLMs), which have significantly advanced the state-of-the-art across a wide range of tasks, including Question Answering (QA). Despite these advancements, research on LLMs has primarily targeted high-resourced languages (e.g., English), and only recently has attention shifted toward multilingual models. However, these models demonstrate a training data bias towards a small number of popular languages or rely on transfer learning from high- to under-resourced languages; this may lead to a misrepresentation of social, cultural, and historical aspects. To address this challenge, monolingual LLMs have been developed for under-resourced languages; however, their effectiveness remains less studied when compared to multilingual counterparts on language-specific tasks. In this study, we address this research gap in Greek QA by contributing: (i) DemosQA, a novel dataset, which is constructed using social media user questions and community-reviewed answers to better capture the Greek social and cultural zeitgeist; (ii) a memory-efficient LLM evaluation framework adaptable to diverse QA datasets and languages; and (iii) an extensive evaluation of 11 monolingual and multilingual LLMs on 6 human-curated Greek QA datasets using 3 different prompting strategies. We release our code and data to facilitate reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。