构建首个口语化教学反馈数据集,提升AI对学习者友好型纠错能力。
Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation

- 基于口语转录文本生成可操作、适龄且鼓励性的教学反馈。
- 监督微调比偏好对齐在纠错与反馈上表现更优,两者质量关联弱。
- 适合教育AI、语言学习系统研发者参考,推动智能辅导发展。
语法错误纠正(GEC)与解释(GEE)已取得显著进展,但真实教学场景还需具备可操作性、适龄性和鼓励性的学习者友好型教学反馈。本文提出SPFG(Spoken Pedagogical Feedback Generation)数据集,基于Speak & Improve Challenge 2025语料库,将流畅性导向的转录文本与GEC目标配对,并包含经人工验证的教师风格反馈,包括优选/次选反馈对以支持偏好学习。研究了基于转录文本的口语化语法错误纠正(SGEC)设定,评估了三款指令微调大模型(Qwen2.5、Llama-3.1、GLM-4),比较监督微调(SFT)与基于偏好对齐的方法(DPO和KTO)在联合生成纠正结果与反馈上的效果。结果显示,SFT带来最一致的提升,而DPO/KTO仅产生较小或混合收益,且纠正质量与反馈质量呈弱耦合关系。代码已开源:https://github.com/Skywalker-Harrison/spfg。
原文摘要 · Abstract (English)
Grammatical error correction (GEC) and explanation (GEE) have made rapid progress, but real teaching scenarios also require \emph{learner-friendly pedagogical feedback} that is actionable, level-appropriate, and encouraging. We introduce \textbf{SPFG} (\textbf{S}poken \textbf{P}edagogical \textbf{F}eedback \textbf{G}eneration), a dataset built based on the Speak \& Improve Challenge 2025 corpus, pairing fluency-oriented transcriptions with GEC targets and \emph{human-verified} teacher-style feedback, including preferred/rejected feedback pairs for preference learning. We study a transcript-based Spoken Grammatical Error Correction (SGEC) setting and evaluate three instruction-tuned LLMs (Qwen2.5, Llama-3.1, and GLM-4), comparing supervised fine-tuning (SFT) with preference-based alignment (using DPO and KTO) for jointly generating corrections and feedback. Results show that SFT provides the most consistent improvements, while DPO/KTO yield smaller or mixed gains, and that correction quality and feedback quality are weakly coupled. Our implementation is available at https://github.com/Skywalker-Harrison/spfg.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。