arXiv:2603.23682cs.HCcs.AI2026-03被引 1

用心理测量法识别人类与聊天机器人答题差异,助力教育评估防作弊

Assessment Design in the AI Era: A Method for Identifying Items Functioning Differentially for Humans and Chatbots

  • 基于差异项目功能分析,检测题目中人与AI的答题偏差
  • 在化学诊断题和高考题上验证,发现12道题存在显著差异
  • 适合教育评估设计者、考试命题人员参考

大型语言模型(LLMs)在教育领域的快速应用对评估设计提出了深刻挑战。为应对基于LLM工具的使用,必须以可推广、有效且可靠的方式刻画其能力优劣。然而,现有评估多依赖基准测试的描述性统计,缺乏基于理论的测量方法来系统比较人类学习者与LLMs的表现差异,难以直接支持评估设计。本文结合教育数据挖掘与心理测量学理论,提出一种统计严谨的方法,用于识别人类与LLMs在答题上存在系统性差异的题目,揭示评估可能被AI滥用的薄弱环节,并定位使生成式AI表现过强或过弱的任务维度。该方法基于差异项目功能(DIF)分析——传统用于检测群体偏见——并结合负向控制分析与项目-总分相关性判别分析。在高一化学诊断测试和大学入学考试两套工具上,对六款主流聊天机器人(ChatGPT-4o & 5.2, Gemini 1.5 & 3 Pro, Claude 3.5 & 4.5 Sonnet)及真实学生回答进行评估。领域专家进一步分析了被DIF标记的题目,归纳出导致聊天机器人表现异常的任务特征。结果表明,基于DIF的分析框架能有效揭示人类与大模型能力的差异边界,为构建更有效、可靠且公平的智能时代评估体系提供支持。

原文摘要 · Abstract (English)

The rapid adoption of large language models (LLMs) in education raises profound challenges for assessment design. To adapt assessments to the presence of LLM-based tools, it is crucial to characterize the strengths and weaknesses of LLMs in a generalizable, valid and reliable manner. However, current LLM evaluations often rely on descriptive statistics derived from benchmarks, and little research applies theory-grounded measurement methods to characterize LLM capabilities relative to human learners in ways that directly support assessment design. Here, by combining educational data mining and psychometric theory, we introduce a statistically principled approach for identifying items on which humans and LLMs show systematic response differences, pinpointing where assessments may be most vulnerable to AI misuse, and which task dimensions make problems particularly easy or difficult for generative AI. The method is based on Differential Item Functioning (DIF) analysis -- traditionally used to detect bias across demographic groups -- together with negative control analysis and item-total correlation discrimination analysis. It is evaluated on responses from human learners and six leading chatbots (ChatGPT-4o \& 5.2, Gemini 1.5 \& 3 Pro, Claude 3.5 \& 4.5 Sonnet) to two instruments: a high school chemistry diagnostic test and a university entrance exam. Subject-matter experts then analyzed DIF-flagged items to characterize task dimensions associated with chatbot over- or under-performance. Results show that DIF-informed analytics provide a robust framework for understanding where LLM and human capabilities diverge, and highlight their value for improving the design of valid, reliable, and fair assessment in the AI era.

评估设计AI检测心理测量教育测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。