arXiv:2606.01584cs.CLcs.AI2026-06中稿 · AIED 2026

发现大模型在教学对话中高自信地误判性别偏见,可能误导学生。

Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents

论文配图:Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents
图 1 · 摘自论文原文
  • 通过重构师生对话并注入可控偏见,构建自然教学场景下的偏见评估数据集。
  • 顶尖大模型在教学对话中对偏见识别准确率低,且错误判断时信心极高。
  • 高自信错误反馈会误导学习者,适合教育AI安全研究者关注。

对话式辅导代理能提升学习参与度和成效,大语言模型(LLMs)正被广泛用于提供可扩展、个性化的反馈。然而,LLMs可能延续或放大刻板社会偏见,在教育场景中带来特殊风险。本研究评估了LLMs在对话式辅导中的高自信偏见识别能力,即模型在无法识别偏见判断的同时仍保持强烈自信,可能影响其推理与对学生反馈的质量。我们提出一种新的数据生成方法,通过重生成学生-智能导师互动,并引入源自基准数据集的受控偏见话轮,实现自然教学情境下的偏见评估。利用该数据,我们评估了多个LLMs对刻板偏见的检测能力,并通过计算与人工评估分析其回答中的置信度与推理过程。结果表明,相较于基准评估,对话式辅导场景中偏见检测更具挑战性,且当前最先进的LLMs在错误判断刻板偏见陈述时表现出过度自信。此外,模型置信度显著影响其推理与反馈内容,凸显了基于LLM的辅导代理出现高自信偏见行为的风险。最后,讨论了相关影响、缓解策略及未来研究方向。

原文摘要 · Abstract (English)

Conversational tutoring agents have been shown to improve learning engagement and student outcomes, and large language models (LLMs) are increasingly used in these systems to provide scalable, personalized feedback. However, LLMs may perpetuate or amplify stereotypical social biases, posing particular risks in educational settings. In this study, we evaluate LLMs in conversational tutoring scenarios to identify high-confidence social biases, instances where models are unable to identify biased judgments in tutoring conversations while maintaining strong confidence in their assessments, potentially affecting their reasoning and the feedback they provide to learners. We present a new dataset generation method that enables bias evaluation under naturalistic instructional conditions by regenerating student-AI tutor interactions and introducing turns with controlled bias derived from a benchmark dataset. Using this data, we assess multiple LLMs' ability to detect stereotypical biases and analyze the confidence and reasoning underlying their responses through computational and human evaluations. We find that bias detection is substantially more challenging in conversational tutoring contexts than in benchmark-based evaluations, and that state-of-the-art LLMs are overconfident in their incorrect assessments of stereotypical bias statements. Moreover, model confidence strongly influences reasoning and feedback, highlighting the risks of overconfident, biased behavior in LLM-based tutoring agents. We conclude by discussing implications, mitigation considerations, and directions for future research.

大模型安全教育AI偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。