用动态奖励机制提升填空题干扰项生成质量
DualReward: A Dynamic Reinforcement Learning Framework for Cloze Tests Distractor Generation
- 双奖励机制区分人工与模型生成的干扰项
- 在跨领域数据上提升3.48%-3.86%准确率
- 适合自动化试题生成与教育AI开发
本文提出DualReward,一种用于填空题干扰项自动生成的新型强化学习框架。与依赖监督学习或静态生成模型的传统方法不同,该框架采用双奖励结构并引入自适应缩放机制,有效区分人工标注的优质干扰项与模型生成候选项。奖励信号强度随模型表现和置信度动态调整。我们在篇章级(CLOTH-F)和句子级(MCQ)填空测试数据集上评估该方法,结果表明:在同质数据集(CLOTH-F)上获得稳定小幅提升,在多样化的跨域数据(MCQ)上实现显著改进(P@1提升3.48%-3.86%),证明其在处理多种题型与领域时具有更强适应性。该工作提供了一个灵活框架,能有效平衡学习可靠人工样本与探索高质量新干扰项之间的关系。
原文摘要 · Abstract (English)
This paper introduces DualReward, a novel reinforcement learning framework for automatic distractor generation in cloze tests. Unlike conventional approaches that rely primarily on supervised learning or static generative models, our method employs a dual reward structure with adaptive scaling that differentiates between human-created gold standard distractors and model-generated candidates. The framework dynamically adjusts reward signal intensity based on model performance and confidence. We evaluate our approach on both passage-level (CLOTH-F) and sentence-level (MCQ) cloze test datasets, demonstrating consistent improvements over state-of-the-art baselines. Experimental results show that our adaptive reward scaling mechanism provides modest but consistent benefits on homogeneous datasets (CLOTH-F) and more substantial improvements (3.48-3.86% in P@1) on diverse, cross-domain data (MCQ), suggesting its particular effectiveness for handling varied question types and domains. Our work offers a flexible framework that effectively balances learning from reliable human examples while exploring novel, high-quality distractors for automated test generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。