arXiv:2507.00883cs.CL2025-07被引 5

测试数学题文化差异对大模型的影响,发现西方题干更易解

Mathematics Isn't Culture-Free: Probing Cultural Gaps via Entity and Scenario Perturbations

  • 用提示工程生成非洲、中国等五地文化适配的数学题
  • 所有模型在本土化题干上表现下降10%-25%
  • 具备推理能力的模型更能适应文化差异

尽管数学常被视为无文化偏见,但其问题呈现方式隐含文化背景。现有基准如GSM8K主要基于西方范式,包含姓名、货币和日常场景。本文通过提示工程结合人工验证,为非洲、印度、中国、韩国和日本生成了GSM8K的本地化变体。评估了六种参数量从8B到72B的大语言模型,采用五种提示策略,考察其对文化差异的鲁棒性。结果表明:模型在原始美国主导的数据集上表现最佳,在文化适配版本上普遍表现下降10%-25%。但具备推理能力的模型更具韧性,说明深度推理有助于缓解数学任务中的文化呈现差距。

原文摘要 · Abstract (English)

Although mathematics is often considered culturally neutral, the way mathematical problems are presented can carry implicit cultural context. Existing benchmarks like GSM8K are predominantly rooted in Western norms, including names, currencies, and everyday scenarios. In this work, we create culturally adapted variants of the GSM8K test set for five regions Africa, India, China, Korea, and Japan using prompt-based transformations followed by manual verification. We evaluate six large language models (LLMs), ranging from 8B to 72B parameters, across five prompting strategies to assess their robustness to cultural variation in math problem presentation. Our findings reveal a consistent performance gap: models perform best on the original US-centric dataset and comparatively worse on culturally adapted versions. However, models with reasoning capabilities are more resilient to these shifts, suggesting that deeper reasoning helps bridge cultural presentation gaps in mathematical tasks

数学推理文化偏见大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。