用轻量模型自动批改孟加拉语答题,更准更省资源。
Semantic Grading of Written Answers in Low-Resource Language Bangla Using a Fine-Tuned Lightweight Language Model

- 微调轻量模型,比对题干、参考答案与学生作答的语义
- 在人工评估中相关性达0.936,误差仅0.725分
- 适合农村缺教师的教育场景,反馈不易泄露答案
孟加拉语是全球使用最广泛的语言之一,但在教育自然语言处理研究中仍被严重忽视。在许多偏远和农村地区,合格学科教师稀缺,书面答案主要依赖人工批改,导致反馈延迟且不一致。自动评分面临挑战,因为语义正确的回答在表达形式上差异很大。我们提出一种双语(孟加拉语-英语)评估系统,专为低资源教育环境设计,优先关注语义正确性而非词汇重合度。该方法微调一个轻量级语言模型,基于题干、参考答案和学生作答生成评分与简明、上下文相关的反馈,适用于课堂部署。我们还构建了一个合成双语数据集以支持可控训练与评估。在统一协议下对比多种专有与开源大模型,经QLoRA微调的Qwen3-8B在合成评估中表现最佳,反馈泄漏率最低(RoRa = 0.819),在专门的人工评估中与人类评分一致性最强(rho = 0.936,MAE = 0.725)。
原文摘要 · Abstract (English)
Bangla is among the world's most widely spoken languages, yet it remains underserved in educational NLP research. In many remote and rural regions, access to qualified subject teachers is limited, and written answers are consequently graded largely by hand, restricting timely and consistent feedback. Automatic assessment is challenging because semantically correct responses can vary substantially in surface form. We present a bilingual (Bangla-English) evaluation system designed for low-resource educational settings that prioritizes semantic correctness over lexical overlap. Our approach fine-tunes a lightweight language model to grade each response using the question, reference answer, and student answer, producing a numeric score and concise, context-grounded feedback suitable for classroom deployment. We also construct a synthetic bilingual dataset to enable controlled training and evaluation. Across proprietary and open-source LLMs evaluated under a unified protocol, our QLoRA-tuned Qwen3-8B confirms consistent improvement by producing the most leakage-resistant feedback (RoRa = 0.819) in synthetic evaluation and the strongest agreement with human scores (rho = 0.936, MAE = 0.725) in a dedicated human study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。