用少量保密数据生成大规模评分训练集,提升自动作文评分效果。
Training data generation for context-dependent rubric-based short answer grading
- 基于小规模保密数据,通过简单文本格式转换生成大量替代数据
- 生成的三个模拟数据集在表面特征上更接近原始数据
- 初步实验显示新方法可能提升自动评分模型性能
每四年,经合组织(OECD)会发布PISA测试,评估全球青少年学生的知识水平并比较各国教育体系。但由于需避免语言差异和评阅者偏见,学生作答的评分极具挑战性。因此,探索自动评分方法颇具意义。训练这类方法(尤其是机器学习模型)或调整非学习类方法的参数与超参,均需大量领域特定数据。本文探索了几种仅依赖少量保密数据作为参考,生成大规模训练数据集的方法,利用极简的文本格式转换策略以保障数据隐私。通过这些方法,我们成功构建了三个替代数据集,其表面特征至少比直接提示生成的结果更贴近原始参考数据。初步实验表明,其中一种方法可能有助于提升自动评分模型的训练效果。
原文摘要 · Abstract (English)
Every four years, the PISA test is administered by the OECD to test the knowledge of teenage students worldwide and allow for comparisons of educational systems. However, having to avoid language differences and annotator bias makes the grading of student answers challenging. For these reasons, it would be interesting to consider methods of automatic student answer grading. To train some of these methods, which require machine learning, or to compute parameters or select hyperparameters for those that do not, a large amount of domain-specific data is needed. In this work, we explore a small number of methods for creating a large-scale training dataset using only a relatively small confidential dataset as a reference, leveraging a set of very simple derived text formats to preserve confidentiality. Using the proposed methods, we successfully created three surrogate datasets that are, at the very least, superficially more similar to the reference dataset than a straightforward result of prompt-based generation. Early experiments suggest one of these approaches might also lead to improved training of automatic answer grading models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。