arXiv:2603.12471cs.CLcs.HC2026-03被引 2

研究发现,自动作文反馈系统会因学生背景差异给出不同偏见性评价。

Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing Feedback

  • 用四个大模型分析600篇作文,测试不同学生属性下的反馈差异。
  • 相同内容下,少数族裔、残障等学生获更多表扬但更少实质性批评。
  • 揭示算法反馈存在隐性偏见,需透明化与问责机制。

有效的个性化反馈对学生的读写能力发展至关重要。尽管当前基于大语言模型(LLM)的工具可规模化提供自动化反馈,但这些模型并非语言中立:它们偏好标准学术英语并复制社会刻板印象,引发人们对“个性化”如何影响反馈质量的担忧。本研究考察了四种主流大模型(GPT-4o、GPT-3.5-turbo、Llama-3.3 70B、Llama-3.1 8B)在学生属性(性别、种族/族裔、学习需求、学业表现、动机)提示下的反馈调整行为。基于来自PERSUADE数据集的600篇八年级论说文,我们生成了嵌入不同属性信息的反馈,并通过改编的Marked Words框架分析词汇变化。结果表明,即使文章内容完全相同,模型仍会根据预设的学生属性系统性地产生符合刻板印象的反馈差异——针对被标记为种族、语言或残障的学生,普遍存在正面反馈偏差(过度表扬)和反馈回避偏差(缺乏实质性批评,假设能力有限)。各属性下,模型不仅调整强调内容,还改变了评判标准与称呼方式。我们将其称为‘标记教学法’(Marked Pedagogies),并呼吁提升自动化反馈工具的透明度与问责机制。

原文摘要 · Abstract (English)

Effective personalized feedback is critical to students' literacy development. Though LLM-powered tools now promise to automate such feedback at scale, LLMs are not language-neutral: they privilege standard academic English and reproduce social stereotypes, raising concerns about how "personalization" shapes the feedback students receive. We examine how four widely used LLMs (GPT-4o, GPT-3.5-turbo, Llama-3.3 70B, Llama-3.1 8B) adapt written feedback in response to student attributes. Using 600 eighth-grade persuasive essays from the PERSUADE dataset, we generated feedback under prompt conditions embedding gender, race/ethnicity, learning needs, achievement, and motivation. We analyze lexical shifts across model outputs by adapting the Marked Words framework. Our results reveal systematic, stereotype-aligned shifts in feedback conditioned on presumed student attributes--even when essay content was identical. Feedback for students marked by race, language, or disability often exhibited positive feedback bias and feedback withholding bias--overuse of praise, less substantive critique, and assumptions of limited ability. Across attributes, models tailored not only what content was emphasized but also how writing was judged and how students were addressed. We term these instructional orientations Marked Pedagogies and highlight the need for transparency and accountability in automated feedback tools.

AI教育语言偏见个性化反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。