用20个问题猜谜游戏发现大模型对全球南方实体推理能力更差
The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities
- 让大模型主动提问,通过多轮推理游戏测试其地理实体推断能力
- 模型对全球北方和西方的实体识别准确率显著高于全球南方和东方
- 语言种类影响小,但预训练数据分布无法完全解释推理差异
大型语言模型虽经调优以减少显性偏见,但仍可能隐含源于预训练数据的细微偏见。我们不使用人工设计的问题触发防护机制,而是研究模型主动提问时的行为表现。将20个问题游戏作为多轮推理任务的测试平台,我们基于新构建的数据集Geo20Q+系统评估了不同地理区域的实体推断性能差异,该数据集包含来自多样地区的人物与文化标志性对象(如食物、地标、动物)。我们在两种游戏配置(标准20轮与无限轮)下测试了多种主流大模型,并覆盖英语、印地语、中文、日语、法语、西班牙语和土耳其语共七种语言。结果揭示显著地理差距:模型对全球北方的实体推断成功率远高于全球南方,对全球西部分辨能力也优于全球东部分。尽管维基百科浏览量和预训练语料频率与性能存在微弱相关性,但无法完全解释这些差异。值得注意的是,游戏语言对性能差距影响甚微。这些发现表明,创造性的自由形式评估框架有助于揭示传统提示设置下隐藏的细微偏见。通过分析模型在多轮中如何发起并推进推理目标,我们发现其推理过程嵌入了地理与文化差异。数据集(Geo20Q+)与代码已公开于 https://sites.google.com/view/llmbias20q/home。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been extensively tuned to mitigate explicit biases, yet they often exhibit subtle implicit biases rooted in their pre-training data. Rather than directly probing LLMs with human-crafted questions that may trigger guardrails, we propose studying how models behave when they proactively ask questions themselves. The 20 Questions game, a multi-turn deduction task, serves as an ideal testbed for this purpose. We systematically evaluate geographic performance disparities in entity deduction using a new dataset, Geo20Q+, consisting of both notable people and culturally significant objects (e.g., foods, landmarks, animals) from diverse regions. We test popular LLMs across two gameplay configurations (canonical 20-question and unlimited turns) and in seven languages (English, Hindi, Mandarin, Japanese, French, Spanish, and Turkish). Our results reveal geographic disparities: LLMs are substantially more successful at deducing entities from the Global North than the Global South, and the Global West than the Global East. While Wikipedia pageviews and pre-training corpus frequency correlate mildly with performance, they fail to fully explain these disparities. Notably, the language in which the game is played has minimal impact on performance gaps. These findings demonstrate the value of creative, free-form evaluation frameworks for uncovering subtle biases in LLMs that remain hidden in standard prompting setups. By analyzing how models initiate and pursue reasoning goals over multiple turns, we find geographic and cultural disparities embedded in their reasoning processes. We release the dataset (Geo20Q+) and code at https://sites.google.com/view/llmbias20q/home.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。