arXiv:2608.28329cs.CLcs.AI2026-08中稿 · and presented at 3…

为孟加拉语医学领域构建了首个问答系统,解决语言资源匮乏问题。

BanglaMed-QA: A Question Answering System for Healthcare Support in Bangla

  • 基于4493个问答对构建结构化医疗知识库,涵盖506种疾病
  • 采用SVM模型实现95% F1分数的问句分类,人类满意度达0.9/1.0
  • 适合低资源语言医疗信息获取者及本土化健康服务开发者

医疗问答系统已成为提供可靠健康信息的重要工具,但针对孟加拉语等低资源语言的研究仍十分有限,主要受限于数据集和专用系统稀缺。为此,本文提出BanglaMed-QA,一个专为孟加拉语医学领域设计的稳健问答系统。该系统首先构建包含4,493个问答对、覆盖9个类别、涉及506种疾病的结构化医疗知识库。为增强语义理解,引入领域特异性词根词典与同义词集,并结合词性标注实现指代消解。采用监督学习模型,其中支持向量机(SVM)在问句分类任务中表现最佳。通过余弦、杰卡德、BM25和莱文斯坦等多种相似度度量,结合软硬投票策略进行查询匹配。系统在自动评估中达到95%的F1分数,在人工评估中平均满意度达0.9/1.0。结果验证了BanglaMed-QA在弥合孟加拉语使用者医疗信息差距方面的实际应用潜力。

原文摘要 · Abstract (English)

Medical question answering (QA) systems have become crucial tools for providing reliable health information. But they remain very unexplored for low-resource languages like Bangla due to limited datasets and systems tailored to these languages. To address this, we introduce BanglaMed-QA, a robust QA system specifically designed for the Bangla medical domain. The process begins with building a structured medical knowledge base that includes 4,493 QA pairs in 9 categories under 506 diseases. To improve semantic comprehension, domain-specific root word dictionaries and synonym sets are proposed, in addition to part-of-speech tagging for anaphora resolution. We adopt supervised machine learning models in which SVM is found to be the best model to categorize questions. Multiple similarity metrics, including cosine, Jaccard, BM25, and Levenshtein, are applied with soft and hard voting methods for query matching. The performance of the QA system has been evaluated in two aspects, with a 95% F1 score in an automated evaluation and an average human satisfaction rating of 0.9 out of 1.0. This validates the real-world application of BanglaMed-QA in closing the healthcare information gap for Bangla speakers.

医疗问答低资源语言孟加拉语知识库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。