arXiv:2509.17701cs.CLcs.AI2025-09被引 5

测试多语言大模型解数学题能力,发现英语答案评分最高,阿拉伯语常被低估。

Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs

  • 构建自动多语言数学题生成与评估流水线,覆盖德、英、阿三语
  • 628道题目测试显示英语解答平均得分显著高于阿拉伯语
  • 用多个大模型作答+人工评判,揭示教育AI中的语言偏见问题

大型语言模型(LLMs)在教育支持中应用日益广泛,但其回答质量受交互语言影响。本文提出一个自动化多语言流水线,用于生成、求解和评估符合德国中小学课程标准的数学题。共生成628道数学练习题,并翻译为英语、德语和阿拉伯语。使用三种商用大模型(GPT-4o-mini、Gemini 2.5 Flash、Qwen-plus)在每种语言中生成分步解答。通过包含Claude 3.5 Haiku在内的多模型评委组,采用对比评估框架对解答质量进行评判。结果显示,英语解答始终获得最高评分,阿拉伯语解答则普遍偏低。这些发现揭示了持续存在的语言偏见,凸显教育领域实现更公平多语言AI系统的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used for educational support, yet their response quality varies depending on the language of interaction. This paper presents an automated multilingual pipeline for generating, solving, and evaluating math problems aligned with the German K-10 curriculum. We generated 628 math exercises and translated them into English, German, and Arabic. Three commercial LLMs (GPT-4o-mini, Gemini 2.5 Flash, and Qwen-plus) were prompted to produce step-by-step solutions in each language. A held-out panel of LLM judges, including Claude 3.5 Haiku, evaluated solution quality using a comparative framework. Results show a consistent gap, with English solutions consistently rated highest, and Arabic often ranked lower. These findings highlight persistent linguistic bias and the need for more equitable multilingual AI systems in education.

多语言教育AI偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。