用语言模型模拟学习者,直接评估问题对考试成绩的提升效果。
Which Questions Improve Learning the Most? Utility Estimation of Questions with LM-based Simulations
- 通过语言模型模拟学习过程,量化问题对考试表现的直接影响。
- 经QUEST训练生成的问题使模拟考试成绩提升超20%。
- 适合需要高效提升学习效果的教育技术研究者与开发者。
提出QUEST(基于语言模型的模拟测试问题效用评估)框架,利用语言模型模拟学习者在阅读教材章节时提问并获取答案,最终参加章节测验。通过该模拟流程,直接衡量每个问题对考试成绩的实际贡献,而非依赖内容显著性或信息增益等间接指标。为此构建了包含五个学科的TEXTBOOK-EXAM基准数据集,将教材段落与章末测验题对齐。利用QUEST筛选高价值问题,并通过拒绝采样微调问题生成模型。实验表明,经过QUEST训练的模型生成的问题,使模拟测试成绩比使用间接指标或提示方法微调的强基线高出20%以上。此外,问题效用与显著性、与测验题相似性仅弱相关,说明其捕捉到独特且有益于学习结果的信号。该框架为问题评估与生成提供以学习成果为导向的新范式。
原文摘要 · Abstract (English)
Asking good questions is critical for comprehension and learning, yet evaluating and generating such questions remains a challenging problem. Prior work on inquisitive questions focuses on learner-generated, curiosity-driven queries and evaluates them using indirect metrics, such as salience or information gain, that do not directly capture a question's impact on actual learning outcomes. We introduce QUEST (Question Utility Estimation with Simulated Tests), a framework that uses language models to simulate learners and directly quantify the utility of a question - its contribution to exam performance. QUEST simulates a learner who asks questions and receives answers while studying a textbook chapter, then uses them to take an end-of-chapter exam. Through this simulation, the utility of each question is estimated by its direct effect on exam performance, rather than inferred indirectly based on the underlying content. To support this evaluation, we curate TEXTBOOK-EXAM, a benchmark that aligns textbook sections with end-of-section exam questions across five academic disciplines. Using QUEST, we filter for high-utility questions and fine-tune question generators via rejection sampling. Experiments show that questions generated by QUEST-trained models improve simulated test scores by over 20% compared to strong baselines that are fine-tuned using indirect metrics or leverage prompting methods. Furthermore, utility is only weakly correlated with salience and similarity to exam questions, suggesting that it captures unique signal that benefits downstream performance. QUEST offers a new outcome-driven paradigm for question evaluation and generation - one that moves beyond question-answer content toward measurable improvements in learning outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。