构建高质量孟加拉语自然语言推理数据集,解决标注混乱问题
BNLI: A Linguistically-Refined Bengali Dataset for Natural Language Inference
- 通过严格标注流程提升语义清晰度和三类关系平衡
- 多模型测试显示新数据集显著提升推理性能与可解释性
- 适合低资源语言推理研究者使用,推动本土化NLP发展
尽管自然语言推理(NLI)研究进展迅速,但孟加拉语资源仍极为有限。现有孟加拉语NLI数据集存在标注错误、句对模糊及语言多样性不足等问题,严重影响模型训练与评估效果。为此,我们提出BNLI——一个经过语言学精细化处理的孟加拉语NLI数据集,旨在支持更可靠的语义理解与推理建模。该数据集通过严谨的标注流程,确保语义清晰且三类关系(蕴含、矛盾、中立)分布均衡。我们采用多种前沿Transformer架构(包括多语言及专用于孟加拉语的模型)对BNLI进行了基准测试,结果表明,使用该数据集可显著提升模型在复杂语义关系上的捕捉能力,增强模型可靠性与可解释性,为孟加拉语及其他低资源语言的推理研究奠定坚实基础。
原文摘要 · Abstract (English)
Despite the growing progress in Natural Language Inference (NLI) research, resources for the Bengali language remain extremely limited. Existing Bengali NLI datasets exhibit several inconsistencies, including annotation errors, ambiguous sentence pairs, and inadequate linguistic diversity, which hinder effective model training and evaluation. To address these limitations, we introduce BNLI, a refined and linguistically curated Bengali NLI dataset designed to support robust language understanding and inference modeling. The dataset was constructed through a rigorous annotation pipeline emphasizing semantic clarity and balance across entailment, contradiction, and neutrality classes. We benchmarked BNLI using a suite of state-of-the-art transformer-based architectures, including multilingual and Bengali-specific models, to assess their ability to capture complex semantic relations in Bengali text. The experimental findings highlight the improved reliability and interpretability achieved with BNLI, establishing it as a strong foundation for advancing research in Bengali and other low-resource language inference tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。