收集485份带评分标准的多模态考试答案,支持精准评估研究。
Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment

- 构建包含手写、公式、图表等多模态数据的考试答案库。
- 12个基于成果教育的评分量表共47个评估维度,每份答案有完整评分。
- 适合做自动化评分、反馈生成与隐私保护评估的研究者使用。
本文介绍一个多模态考试答案数据集,包含485份来自4所高校415名学生的扫描答卷,由8位教师提供涵盖9个学科的12种题型。每份答案关联随机标识符、学科标签、题目、标准答案、评分准则定义、表现等级描述、各准则得分及总分。12个评分量表共含47项评估指标。扫描件保留真实学术内容,包括手写、印刷体、公式、表格、代码、图像、草图与示意图。不同设备(CamScanner、Adobe Scan、传统扫描仪)引入光照、对比度、方向、压缩与分辨率差异。多样化的书写风格、划改痕迹、修订计算与补充修正进一步增强视觉多样性,用于模型鲁棒性与泛化性研究。数据整理包括异源数据整合、标签与文本标准化、分数验证、标识符与文件名随机化、以及JSON到PDF完整性检查。答案级审计确认:485个唯一标识符、485个唯一文件名、总分与各准则分之和一致,且分数未超量表上限。该数据集可支持量表感知的自动评估、多模态文档理解、逐项反馈生成、分数预测与隐私保护的成果导向教育研究。数据仅限科研用途,经合理申请可向通讯作者获取。
原文摘要 · Abstract (English)
This data article describes a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE) grading metadata. The collection contains 485 answer submissions from 415 consenting students at four academic institutions. Eight faculty contributors supplied examination materials covering nine subjects and 12 distinct question templates. Each answer-level item links a scanned PDF to a randomized identifier, subject label, question, model answer, criterion definitions, performance-level descriptions, criterion marks, and a total mark. The 12 rubrics contain 47 criteria in total. The scans retain realistic academic content, including handwriting, printed text, equations, tables, code, figures, sketches, and diagrams. CamScanner, Adobe Scan, and conventional scanners contributed variation in illumination, contrast, orientation, compression, and resolution. Diverse handwriting, crossed-out work, revised calculations, and inserted corrections add further visual variability for robustness and generalization studies. Preparation involved heterogeneous-source consolidation, label and text standardization, score validation, identifier randomization, filename randomization, and JSON-to-PDF integrity checks. An answer-level audit confirmed 485 unique identifiers, 485 unique PDF filenames, agreement between each total mark and its criterion-mark sum, and scores within the applicable rubric maximum. The data can support rubric-aware automated evaluation, multimodal document understanding, criterion-level feedback, score prediction, and privacy-aware OBE assessment research. Access is restricted to research use and is available from the corresponding author upon reasonable request.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。