对比方言与标准语的维基知识差异,发现大模型在本地信息上表现差。
Information Asymmetry across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA
- 构建跨方言维基问答数据集,对比粤语/德语方言与普通话/巴伐利亚语
- 大模型无法回答仅存在于本地维基的知识题,但提供摘要后性能显著提升
- 本地维基含区域与全球信息,对大模型文化包容性提出挑战
大型语言模型(LLMs)正成为人类获取知识的常见方式,但其覆盖范围与可靠性存在显著差异。尤其对于地方语言变体,常出现本地维基内容未被标准版本收录的情况。然而,目前尚不清楚大模型在信息不对称场景下的表现,特别是对密切相关的语言变体。我们手动构建了一个新型问答(QA)数据集,涵盖仅存在于本地维基页面的知识,对比了普通话-粤语和德语-巴伐利亚语两种情形。实验表明,大模型难以回答仅出现在本地维基中的问题。提供正文导语部分作为上下文可显著提升性能,进一步通过翻译还能获得增益。基于主题、地理标注及分层评估,揭示了本地维基作为区域与全球信息来源的价值。这些发现引发关于大模型包容性与文化覆盖的重要思考。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are becoming a common way for humans to seek knowledge, yet their coverage and reliability vary widely. Especially for local language varieties, there are large asymmetries, e.g., information in local Wikipedia that is absent from the standard variant. However, little is known about how well LLMs perform under such information asymmetry, especially on closely related languages. We manually construct a novel challenge question-answering (QA) dataset that captures knowledge conveyed on a local Wikipedia page, which is absent from their higher-resource counterparts-covering Mandarin Chinese vs. Cantonese and German vs. Bavarian. Our experiments show that LLMs fail to answer questions about information only in local editions of Wikipedia. Providing context from lead sections substantially improves performance, with further gains possible via translation. Our topical, geographic annotations, and stratified evaluations reveal the usefulness of local Wikipedia editions as sources of both regional and global information. These findings raise critical questions about inclusivity and cultural coverage of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。