用能力条件生成模拟学生写作,让大模型更真实地表现不同水平的作文差异。
SWIM: Student Writing Simulation via Proficiency-Conditioned Generation

- 基于写作能力条件生成,实现多维度写作特征模拟。
- 监督微调显著提升各水平写作特征匹配度,强化学习再进一步优化。
- 适合教育评估、智能辅导系统研究者使用。
写作能力体现在内容发展、思想组织、词汇选择和语言运用等多个方面。尽管基于大语言模型的学生写作模拟日益受到关注,但大模型能否真实再现长篇写作中多维度的能力差异仍缺乏探索。本文提出SWIM任务,将学生写作模拟建模为能力条件下的作文生成问题。通过自动作文评分衡量生成内容与目标能力水平的匹配程度,评估了提示工程、监督微调(SFT)和强化学习(RL)三种方法的效果。实验表明,提示工程对能力控制有限,即使使用强健的闭源大模型并结合评分标准策略,也难以准确还原不同能力水平在词汇、语法和组织结构上的差异。而监督微调显著提升了匹配度,强化学习结合新提出的“能力对齐奖励”在所有写作维度和题目上均取得进一步提升。结果表明,显式监督相比仅靠提示能带来更强的能力特征对齐,但真实低水平写作仍难精准复现。
原文摘要 · Abstract (English)
Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplored. In this work, we explore if language models can realistically simulate student writing, and introduce SWIM, a task that formulates Student Writing sIMulation as proficiency-conditioned essay generation. We evaluate prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) methods for writing simulation using automated essay scoring as a measure of profile alignment. Extensive experiments reveal that prompting provides limited proficiency control, even for strong proprietary LLMs with rubric-grounded strategies. In particular, while models can adjust content-oriented traits, they struggle to reproduce the lexical, grammatical, and organizational variation in different proficiency levels. SFT substantially improves alignment, while RL with the proposed proficiency-alignment reward yields further gains across all writing traits and essay prompts. Our findings suggest that explicit supervision enables substantially stronger profile alignment than prompting alone, while authentic low-proficiency writing remains challenging to reproduce.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。