arXiv:2412.18440cs.CL2024-12被引 1

用BERT模型自动评估孟加拉语教材问答,发现本地化模型效果最佳。

Unlocking the Potential of Multiple BERT Models for Bangla Question Answering in NCTB Textbooks

  • 对比三种BERT模型在孟加拉语教材问答上的表现
  • 本地BERT模型达F1 0.75、EM 0.53,优于其他模型
  • 适合教育科技研究者与多语言NLP开发者参考

评估教育场景中的文本理解能力对了解学生表现和优化课程有效性至关重要。本研究考察了RoBERTa Base、Bangla-BERT和BERT Base三种先进语言模型,在6至10年级国家课程与教科书委员会(NCTB)教材中的孟加拉语篇章问答自动评估能力。构建了约3,000个孟加拉语篇章问答数据集,通过F1分数和精确匹配(EM)指标,在不同超参数配置下评估模型性能。结果表明,Bangla-BERT始终表现最优,最高获得F1 0.75、EM 0.53,尤其在小批量、包含停用词及中等学习率下表现更佳。相反,RoBERTa Base在某些配置下表现最差,最低达F1 0.19、EM 0.27。研究强调了超参数调优对模型性能的重要性,并展示了机器学习模型在教育文本理解评估中的潜力。然而,数据集规模、拼写不一致和计算资源限制仍需进一步研究以提升模型鲁棒性与适用性。本研究为未来教育机构自动化评估系统的发展奠定了基础。

原文摘要 · Abstract (English)

Evaluating text comprehension in educational settings is critical for understanding student performance and improving curricular effectiveness. This study investigates the capability of state-of-the-art language models-RoBERTa Base, Bangla-BERT, and BERT Base-in automatically assessing Bangla passage-based question-answering from the National Curriculum and Textbook Board (NCTB) textbooks for classes 6-10. A dataset of approximately 3,000 Bangla passage-based question-answering instances was compiled, and the models were evaluated using F1 Score and Exact Match (EM) metrics across various hyperparameter configurations. Our findings revealed that Bangla-BERT consistently outperformed the other models, achieving the highest F1 (0.75) and EM (0.53) scores, particularly with smaller batch sizes, the inclusion of stop words, and a moderate learning rate. In contrast, RoBERTa Base demonstrated the weakest performance, with the lowest F1 (0.19) and EM (0.27) scores under certain configurations. The results underscore the importance of fine-tuning hyperparameters for optimizing model performance and highlight the potential of machine learning models in evaluating text comprehension in educational contexts. However, limitations such as dataset size, spelling inconsistencies, and computational constraints emphasize the need for further research to enhance the robustness and applicability of these models. This study lays the groundwork for the future development of automated evaluation systems in educational institutions, providing critical insights into model performance in the context of Bangla text comprehension.

孟加拉语问答系统BERT教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。