用大模型生成真实作文,提升自动评分数据量
Calibrating Generative AI to Produce Realistic Essays for Data Augmentation
- 采用三种提示策略生成学生作文模拟样本
- 预测下一个词策略评分一致性最高,文本最真实
- 适合用于教育AI训练数据增强,提升评估可靠性
数据增强可缓解机器学习自动评分系统在开放作答题中因训练数据不足带来的问题。本研究探讨三种大型语言模型提示方法生成的作文在保留原作文质量及生成文本真实性方面的表现。我们构建了学生作文的模拟版本,由人工评分员对其评分并评估文本真实性。结果显示,预测下一个词提示策略在人工评分者间对模拟作文得分的一致性最高;预测下一个词和句子级提示策略最能保持模拟作文的原始质量评分;而预测下一个词与25个示例提示策略生成的文本在人类评判下最为真实。
原文摘要 · Abstract (English)
Data augmentation can mitigate limited training data in machine-learning automated scoring engines for constructed response items. This study seeks to determine how well three approaches to large language model prompting produce essays that preserve the writing quality of the original essays and produce realistic text for augmenting ASE training datasets. We created simulated versions of student essays, and human raters assigned scores to them and rated the realism of the generated text. The results of the study indicate that the predict next prompting strategy produces the highest level of agreement between human raters regarding simulated essay scores, predict next and sentence strategies best preserve the rated quality of the original essay in the simulated essays, and predict next and 25 examples strategies produce the most realistic text as judged by human raters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。