arXiv:2606.11009cs.CLcs.CY2026-06

评测大模型跨文化数学题翻译,发现模型偏好表面符号而丢失文化多样性。

Who Brought Easter Eggs to Eid? Auditing Cultural Translation of Math Word Problems Across Diverse Languages and Regions

论文配图:Who Brought Easter Eggs to Eid? Auditing Cultural Translation of Math Word Problems Across Diverse Languages and Regions
图 1 · 摘自论文原文
  • 分析3大模型在7种语言间翻译数学题的实体转换模式。
  • 仅33.5%替换结果一致,文化多样性显著萎缩。
  • 适合关注AI公平性与教育本地化的研究者阅读。

大型语言模型被用于大规模个性化数学题适配,但其跨模型、跨语言的文化一致性仍不明确。本文分析Claude Opus 4、GPT-4.1和Gemini 2.5 Pro将60道英文数学题翻译成孟加拉语、印地语、旁遮普语(印度)、乌尔都语、信德语(巴基斯坦)、意大利语及西西里语(意大利)的表现,覆盖从高资源语言到低资源语言的完整谱系。人工标注6,489个实体转换,编码是否保留、本地化、泛化、省略或更改名称、食物、地点等。模型在转换类型上达成共识的比例为62.5%,具体替换一致率仅33.5%,表明模型选择直接影响学生接触的文化世界。所有21种语言-模型组合均出现熵压缩现象,适配反而减少文化多样性。模型优先处理姓名、食物、货币等表层特征,而保留如年级制度等嵌入文化假设的深层结构。尽管提示指定目标国家,模型仍错误使用孟加拉塔卡对应印度孟加拉语学习者,并产生跨文化混杂,如将复活节彩蛋活动改编为开斋节活动。部分错误可从单例翻译中观察,但多样性萎缩、表层特征偏好、系统性区域误判等深层问题仅在语料级分析中显现。表面合理性使错误更易被忽略。

原文摘要 · Abstract (English)

Large language models are increasingly used to adapt math word problems for personalized learning at scale, but it remains an open question whether those adaptations are consistent across models, preserve cultural diversity at scale, and reveal which cultural entities models treat as most salient. We analyze how Claude Opus 4, GPT-4.1, and Gemini 2.5 Pro adapt 60 English math word problems into Bengali, Hindi, Punjabi (India), Urdu, Sindhi (Pakistan), Italian, and Sicilian (Italy), a language set spanning the full resource spectrum, from high-resource Italian and Hindi to under-studied Sindhi, Sicilian, and Punjabi. We annotate 6,489 entity transformations, coding whether models preserve, localize, generalize, omit, or change entities such as names, foods, and places. Models agree on transformation type in 62.5% of cases and on specific substitutions in only 33.5%, meaning model choice directly shapes which cultural world students encounter. All 21 language-model combinations show entropy collapse, with adaptation compressing rather than expanding cultural diversity. Models prioritize surface markers such as names, foods, and currencies while preserving deeper structural features such as grade-level systems that embed culturally specific assumptions. Despite prompts specifying target countries, models misattribute regional context by using Bangladeshi taka for Indian Bengali students and produce cross-cultural contamination, such as adapting egg hunts as Eid activities. Some failures are visible in individual translations. Others, including diversity collapse, systematic preference for surface markers, and consistent regional misattribution, emerge only through corpus-level analysis. The surface plausibility that makes adapted problems look correct is precisely what makes deeper failures easy to overlook.

文化翻译数学题生成多语言模型教育公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。