构建首个平衡的孟加拉语教育问答数据集,提升模型对无答案问题的判断能力。
NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance
- 从50本教材提取8.7万条问答,平衡有答案与无答案题
- BERT经微调后F1提升313%(0.150→0.620),答案质量显著提高
- 适合低资源语言问答研究者,尤其关注教育场景与鲁棒性
低资源语言的阅读理解系统在处理无答案问题时表现不佳,常生成不可靠回答。为解决此问题,我们提出NCTB-QA,一个大规模孟加拉语教育问答数据集,包含从孟加拉国国家课程与教科书委员会出版的50本教材中提取的87,805个问答对。与现有孟加拉语数据集不同,NCTB-QA保持可回答(57.25%)与不可回答(42.75%)问题的均衡分布,并包含带有合理干扰项的对抗性样本。我们对三种基于Transformer的模型(BERT、RoBERTa、ELECTRA)进行基准测试,通过微调实现显著性能提升:BERT的F1分数相对提高313%(从0.150升至0.620)。所有模型在BERTScore衡量的语义答案质量上均有明显改善。结果表明,领域特定微调对低资源环境下模型的鲁棒性至关重要,确立了NCTB-QA作为孟加拉语教育问答的挑战性基准。
原文摘要 · Abstract (English)
Reading comprehension systems for low-resource languages face significant challenges in handling unanswerable questions. These systems tend to produce unreliable responses when correct answers are absent from context. To solve this problem, we introduce NCTB-QA, a large-scale Bangla question answering dataset comprising 87,805 question-answer pairs extracted from 50 textbooks published by Bangladesh's National Curriculum and Textbook Board. Unlike existing Bangla datasets, NCTB-QA maintains a balanced distribution of answerable (57.25%) and unanswerable (42.75%) questions. NCTB-QA also includes adversarially designed instances containing plausible distractors. We benchmark three transformer-based models (BERT, RoBERTa, ELECTRA) and demonstrate substantial improvements through fine-tuning. BERT achieves 313% relative improvement in F1 score (0.150 to 0.620). Semantic answer quality measured by BERTScore also increases significantly across all models. Our results establish NCTB-QA as a challenging benchmark for Bangla educational question answering. This study demonstrates that domain-specific fine-tuning is critical for robust performance in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。