arXiv:2606.04846cs.CL2026-06

测试大模型对美国各州历史课纲的适配度,发现其回应受州政治倾向影响而非真实课纲。

Large Language Models in K-12 Education: Alignment with State Curriculum Standards and Student Personas

论文配图:Large Language Models in K-12 Education: Alignment with State Curriculum Standards and Student Personas
图 1 · 摘自论文原文
  • 用大模型分析各州历史课纲差异,评估模型响应是否匹配实际教学内容。
  • 模型能根据年级调整内容,但对种族性别敏感度低,对州政治倾向反应明显。
  • 提醒教育者警惕开放大模型导致的学习偏差,需加强与课纲对齐的技术。

随着大语言模型(LLMs)在教育场景中日益普及,其使用带来的伦理问题备受关注。公开在线聊天机器人能力快速提升,被学生广泛用于作业求助。由于美国各州课程标准在内容、重点和叙事上存在显著差异,本研究开发了一套基于大模型的流程,识别各州历史课程标准的差异,并评估不同大模型对这些州级课程差异的反映程度。此外,通过控制实验,改变用户属性如地理位置、年级、性别和种族,考察大模型响应对用户特征的敏感性。结果表明,尽管模型能够调整历史话题的呈现方式,但这种调整更多源于对各州政治倾向的感知,而非真实课程内容。同时,模型能有效适应学生年级水平,对种族和性别表现出较低敏感性,说明其具备一定的个性化适应能力且具有有限的群体偏见。综合来看,这些发现揭示了开放访问大模型可能对学生学习成果造成的风险,主要源于与州级课程标准的错位,亟需更稳健的对齐技术。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become increasingly popular in educational settings, they raise important questions about the ethical implications of their use. Publicly available online chatbots are quickly improving in capability and accuracy leading to more widespread use, including among students looking for help with their homework. This makes it crucial to consider whether these models are aligned with educational standards. Because curriculum standards in the United States are set at the state level, they differ significantly in required content, emphasis, and narrative focus. In this work, we develop an LLM-based pipeline to identify variations in U.S. History curricula across states and evaluate the extent to which different LLMs reflect these state-specific curricular differences. In addition, we conduct controlled experiments that vary user personas by stating user attributes such as geographic location, grade level, gender and race to evaluate the sensitivity of LLM responses to user characteristics. We find that while models are able to adjust their presentation of historical topics, these shifts may come from the perceived political leanings of states and do not necessarily reflect actual curriculum content. Additionally, models successfully adapt to a student's grade level while showing minimal sensitivity to race or gender, suggesting they are capable of useful adaptation to student personas with limited demographic bias. Together, these findings highlight potential risks that open access to LLM chatbots may cause to student learning outcomes stemming from misalignment with state curriculum standards and highlight the need for more robust alignment techniques.

大模型教育应用课纲对齐个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。