用微调的BERT-Bangla构建孟加拉语闭域问答系统
Question-Answering System for Bangla: Fine-tuning BERT-Bangla for a Closed Domain
- 基于孟加拉语BERT模型,针对特定领域微调训练
- 在2500组问答对上达到55.26%精确匹配率
- 适合关注小语种NLP与本地化问答应用的研究者
孟加拉语问答系统的发展相对有限,尤其在特定领域应用方面。本文利用自然语言处理进展,提出一个基于微调BERT-Bangla模型的闭域问答系统。数据源自库尔纳工程技术大学(KUET)网站及其他相关文本,共构建2500组问答对用于训练与评估。采用精确匹配(EM)和F1分数作为关键评价指标,分别取得55.26%和74.21%的成绩。结果表明该系统在特定领域中具有良好潜力,未来仍需优化以应对更复杂查询。
原文摘要 · Abstract (English)
Question-answering systems for Bengali have seen limited development, particularly in domain-specific applications. Leveraging advancements in natural language processing, this paper explores a fine-tuned BERT-Bangla model to address this gap. It presents the development of a question-answering system for Bengali using a fine-tuned BERT-Bangla model in a closed domain. The dataset was sourced from Khulna University of Engineering \& Technology's (KUET) website and other relevant texts. The system was trained and evaluated with 2500 question-answer pairs generated from curated data. Key metrics, including the Exact Match (EM) score and F1 score, were used for evaluation, achieving scores of 55.26\% and 74.21\%, respectively. The results demonstrate promising potential for domain-specific Bengali question-answering systems. Further refinements are needed to improve performance for more complex queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。