测试大模型能否像有误解的学生一样互动,发现它们更像盲目服从的答题者。
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators

- 设计对比反馈机制,判断模型是否只在针对错误认知时修正答案
- 7个模型表现接近零分,无论反馈是否相关都几乎同样改错
- 提出新训练方法,让模型学会基于认知误解持续调整回答
大语言模型能生成类似学生作答的文本,常被用作模拟学生以训练或评估智能辅导系统。然而现有评估多依赖输出与真实学生相似度,而非其是否表现出连贯的认知误解。本文提出一种可控评估框架,衡量模拟器是否在交互中保持误解驱动的信念状态,并仅在反馈针对核心误解时修正答案。核心是误解对比反馈协议:将针对性反馈与不匹配反馈(针对其他合理误解)和通用反馈(仅指出答案错误)对比。提出选择性翻转分数(SFS),量化模型在针对性反馈下更频繁改错的程度。在7个不同规模(4B-120B)的LLM、多个数据集和提示策略下,所有模型SFS接近零,表明其纠错行为不受反馈相关性影响。进一步分析揭示‘谄媚式失败’模式:模型不似持固定误解的学生,反而像对任何纠正信号都立即放弃原观点、从内部知识重新求解的问题解决者。为此,我们开发了结合监督微调(SFT)、偏好优化与基于SFS对齐奖励的强化学习(RL)的后训练流程;SFT带来显著提升(最高+0.56),而SFS对齐的RL比偏好优化更具一致性。结果表明,认知误解忠实性是可训练的挑战性属性,推动学生建模从静态输出匹配转向动态信念感知。
原文摘要 · Abstract (English)
Large language models (LLMs) can fluently generate student-like responses, making them attractive as simulated students for training and evaluating AI tutors and human educators. Yet such simulators are typically evaluated by output similarity to real students, not by whether they behave like students with coherent misconceptions during interaction. We introduce a controlled framework for evaluating misconception faithfulness, whether a simulator maintains a misconception-driven belief state and updates selectively when feedback addresses the underlying misconception. Central to our framework is a misconception-contrastive feedback protocol that compares targeted feedback against two controls: misaligned feedback (targeting a different but plausible misconception) and generic feedback (only identifying answer is wrong). We propose Selective Flip Score (SFS), which quantifies how much more often a simulator flips its answer under targeted feedback than under contrastive controls. Across seven LLMs (4B-120B), multiple datasets, and prompting strategies, simulators exhibit near-zero SFS, correcting their answers at similarly high rates regardless of feedback relevance. Further analyses reveal a sycophantic failure mode: models behave less like students with misconceptions but more like problem-solvers who treat any corrective signal as a cue to abandon the simulated belief and re-solve from internal knowledge. To address this, we develop a post-training pipeline spanning supervised fine-tuning (SFT), preference optimization, and reinforcement learning (RL) with an SFS-aligned reward; SFT yields notable gains up to +0.56, and SFS-aligned RL provides more consistent improvements than preference optimization. Our results establish misconception faithfulness as a challenging yet trainable property, motivating a shift from static output matching toward interactive, belief-aware student modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。