arXiv:2604.27618cs.AIcs.CY2026-04

构建数学教育数字影子数据集,揭示大模型在解题中的态度与偏见。

Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs

论文配图:Math Education Digital Shadows for Investigating Learning with GenAI: Mathematics Performance, Anxiety, and Confidence in LLMs
图 1 · 摘自论文原文
  • 设计人类与AI双角色模拟,生成14个大模型的2.8万条数学推理记录
  • 发现大模型在人设模式下出现数学焦虑、逻辑错误和过度自信等现象
  • 适合教育研究者与安全数学助教开发者使用

理解大语言模型(LLMs)对数学教育的影响需要关于其数学表现与偏见的数据。为此,我们提出数学教育数字影子(MEDS)数据集,该数据集通过人类与人工智能类比两种角色,映射大模型在数学推理中的表现。MEDS包含14个大模型(如Mistral、Qwen、DeepSeek、IBM Granite、Microsoft Phi、xAI Grok)在人类影子和AI助手模式下的28,000次运行记录。每条记录包含一组提示、心理与社会人口学元数据,以及四项任务:(i) 关于数学关系的访谈,(ii) 三项测量数学自我效能与焦虑的心理量表,(iii) 一个反映数学态度的认知网络,(iv) 18道高中数学测验题,附带推理解释与信心评分。数据分析显示,大模型在人类影子模式下表现出类似人类的负面数学态度、逻辑谬误及数学上的过度自信。作为数据资源,MEDS可为学习科学与更安全数学助教的开发者提供支持。

原文摘要 · Abstract (English)

Understanding the impact of large language models (LLMs) on mathematics education requires data on LLMs' mathematical performance and biases. To this end, we introduce Math Education Digital Shadows (MEDS), a dataset mapping how LLMs reason about mathematics across human- and AI-like personifications. MEDS comprises 28,000 runs from 14 LLMs (i.e., Mistral, Qwen, DeepSeek, IBM Granite, Microsoft Phi, and xAI Grok) generated under human-shadow and AI-assistant conditions. Each record (digital shadow) includes a set of prompts; psychological and sociodemographic metadata; and four mathematics tasks: (i) interviews about relationships with mathematics, (ii) three psychometric questionnaires on mathematics self-efficacy and anxiety, (iii) one cognitive network capturing attitudes towards mathematics, and (iv) 18 high-school mathematics quiz items enriched with reasoning explanations and confidence scores. Analyses of the data show that LLMs exhibit differences in attitudes and performance across human-shadow and AI-assistant modes, including human-like negative attitudes towards mathematics, logical fallacies, and overconfidence in mathematics. As a data resource, MEDS can benefit learning scientists and developers of safer AI tutors in mathematics.

大模型数学教育认知偏差数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。