大模型在考试中表现好,但教孩子时却常偏离真实教学目标。
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
- 用学生学习任务对比大模型与人类教师表现
- 模型间偏差相关性高于模型与人类的匹配度,且常与学习效果负相关
- 多模型集成反而加剧偏差,预训练共性是主要问题
大语言模型在各类基准测试中表现优异,但这一表现并不保证其在下游任务中的有效性。本研究对比了主流大模型在中小学教学任务中的表现,这些任务难以验证。结果显示,不同模型间的响应差异相关性远高于其与人类专家在目标任务上的相关性。模型共有的偏差与教学质量和学生实际学习成果的预期影响严重错位,甚至呈负相关。进一步发现,多模型集成(包括一致投票和基于基准表现加权)会加剧这种错位。我们测量到,模型选择与提示策略仅能解释15%的错位误差,而大部分错位在不同模型间共享,表明共同预训练是导致此类错位的主要原因。本文提出了衡量复杂任务对齐性的稳健方法,并为高噪声场景下大模型的实际应用提供了独特洞见。
原文摘要 · Abstract (English)
LLMs increasingly excel on AI benchmarks, but doing so does not guarantee validity for downstream tasks. This study contrasts LLM alignment on benchmarks, downstream tasks, and, importantly the intended impact of those tasks. We evaluate the performance of leading LLMs (i.e., generative pre-trained base models) on difficult-to-verify tasks of the teaching and learning of schoolchildren. Across all LLMs, inter-model behaviors on disparate tasks correlate higher than they do with expert human behaviors on target tasks. These biases shared across LLMs are poorly aligned with downstream measures of teaching quality and often negatively aligned with the intended impact of student learning outcomes. Further, we find multi-model ensembles, both unanimous model voting and expert-weighting by benchmark performance, further exacerbate misalignment with learning. We measure that selection of LLM and/or prompting strategy only reliably accounts for $15\%$ of all measured misalignment error and that variation in misalignment error is shared across LLMs, suggesting that common pretraining accounts for much of the misalignment in these tasks. We demonstrate methods for robustly measuring alignment of complex tasks and provide unique insights into practical applications of LLMs in high-noise contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。