用假设驱动框架让大模型可靠模拟学习者行为。
Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments
- 通过可验证的假设组合构建仿真框架
- GPT-4 Turbo在不同模型下保持行为校准
- 适合教育智能系统开发与评估者
模拟学习者行为有助于压力测试开放式交互式学习环境并提前原型化新适应策略。尽管近期研究显示大语言模型(LLMs)在模拟人类行为方面具有潜力,但此类方法仍停留在初步概念验证阶段,主要受限于两大问题:一是LLM对微小提示变化高度敏感,难以泛化至新场景,需大量提示工程;二是看似成功的输出往往不可靠,可能源于领域专家无意引导模型生成预期结果(自我实现预言),或模型在训练数据中见过高度相似场景,导致其并非模拟行为而是复现记忆内容。为此,我们提出Hyp-Mix仿真创作框架,使专家可通过组合可验证的学习者行为假设来构建和评估仿真。在物理学习环境中测试该框架发现,即使底层学习模型发生变化,GPT-4 Turbo仍能保持校准的行为表现,首次提供证据表明大模型可在开放式交互学习环境中可靠模拟真实学习行为,为实用化的LLM行为模拟奠定基础。
原文摘要 · Abstract (English)
Simulating learner actions helps stress-test open-ended interactive learning environments and prototype new adaptations before deployment. While recent studies show the promise of using large language models (LLMs) for simulating human behavior, such approaches have not gone beyond rudimentary proof-of-concept stages due to key limitations. First, LLMs are highly sensitive to minor prompt variations, raising doubts about their ability to generalize to new scenarios without extensive prompt engineering. Moreover, apparently successful outcomes can often be unreliable, either because domain experts unintentionally guide LLMs to produce expected results, leading to self-fulfilling prophecies; or because the LLM has encountered highly similar scenarios in its training data, meaning that models may not be simulating behavior so much as regurgitating memorized content. To address these challenges, we propose Hyp-Mix, a simulation authoring framework that allows experts to develop and evaluate simulations by combining testable hypotheses about learner behavior. Testing this framework in a physics learning environment, we found that GPT-4 Turbo maintains calibrated behavior even as the underlying learner model changes, providing the first evidence that LLMs can be used to simulate realistic behaviors in open-ended interactive learning environments, a necessary prerequisite for useful LLM behavioral simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。