大模型数学推理受文化背景影响,陌生文化导致准确率下降。
Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?
- 用6个不同国家的文化改编版GSM8K测试14个大模型
- 文化不熟悉时准确率下降0.3%至5.9%,差异显著(p<0.01)
- 训练数据含中东南亚内容的模型在本地文化任务中表现更优
我们证明大型语言模型(LLMs)的数学推理能力具有文化敏感性:在六个文化适配版本的GSM8K基准上测试了来自Anthropic、OpenAI、Google、Meta、DeepSeek、Mistral和Microsoft的14个模型,发现当数学题嵌入陌生文化背景时,即使数学逻辑不变,准确率也下降0.3%(Claude 3.5 Sonnet)至5.9%(LLaMA 3.1-8B)。经McNemar检验确认,该性能下降在统计上显著(p < 0.01)。为生成海地、摩尔多瓦、巴基斯坦、所罗门群岛、索马里和苏里南的适配版本,我们系统替换1,198道GSM8K题目中的文化实体(如人名、食物、地点等),同时保持所有数学运算与数值不变。对18,887个实例的定量错误分析显示,文化适应影响整体推理模式,其中数学推理错误占失败的54.7%,计算错误占34.5%。有趣的是,文化熟悉度可提升表现:因训练数据包含中东和南亚内容,Mistral Saba在巴基斯坦适配问题上优于部分更大模型。本研究强调需更丰富的训练数据以保障全球语境下的鲁棒性。
原文摘要 · Abstract (English)
We demonstrate that large language models' (LLMs) mathematical reasoning is culturally sensitive: testing 14 models from Anthropic, OpenAI, Google, Meta, DeepSeek, Mistral, and Microsoft across six culturally adapted variants of the GSM8K benchmark, we find accuracy drops ranging from 0.3% (Claude 3.5 Sonnet) to 5.9% (LLaMA 3.1-8B) when math problems are embedded in unfamiliar cultural contexts--even when the underlying mathematical logic remains unchanged. These statistically significant performance reductions (p < 0.01, confirmed through McNemar tests) reveal that mathematical reasoning in LLMs is not culturally neutral. To create these variants for Haiti, Moldova, Pakistan, Solomon Islands, Somalia, and Suriname, we systematically replaced cultural entities (names, foods, places, etc.) in 1,198 GSM8K questions while preserving all mathematical operations and numerical values. Our quantitative error analysis of 18,887 instances reveals that cultural adaptation affects broader reasoning patterns, with mathematical reasoning errors comprising 54.7% and calculation errors 34.5% of failures. Interestingly, cultural familiarity can enhance performance: Mistral Saba outperforms some larger models on Pakistan-adapted problems due to Middle Eastern and South Asian training data exposure. This study underscores the need for more diverse training data to ensure robust LLM performance across global contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。