构建首个孟加拉语疾病症状数据集,实现98%准确率的疾病预测。
Bridging the Gap in Bangla Healthcare: Machine Learning Based Disease Prediction Using a Symptoms-Disease Dataset
- 基于758个症状-疾病关系构建孟加拉语医疗数据集
- 集成模型在孟加拉语症状输入下达到98%准确率
- 为孟加拉语人群提供早期疾病诊断工具支持
提升非英语人群获取可靠健康信息的能力至关重要,但孟加拉语疾病预测资源仍十分有限。本研究通过构建包含758个独特症状-疾病关联、覆盖85种疾病的综合性孟加拉语症状-疾病数据集,填补了这一空白。为确保透明性和可复现性,该数据集已公开发布。利用该数据集,我们评估了多种机器学习模型对孟加拉语症状输入的疾病预测性能,并采用软投票与硬投票集成方法,结合表现最佳模型,实现了98%的准确率,展现出优异的鲁棒性和泛化能力。本工作为孟加拉语疾病预测奠定了基础资源,推动本地化健康信息学和诊断工具的发展,助力孟加拉语社区实现更公平的健康信息获取,尤其在早期疾病筛查和医疗干预方面具有重要意义。
原文摘要 · Abstract (English)
Increased access to reliable health information is essential for non-English-speaking populations, yet resources in Bangla for disease prediction remain limited. This study addresses this gap by developing a comprehensive Bangla symptoms-disease dataset containing 758 unique symptom-disease relationships spanning 85 diseases. To ensure transparency and reproducibility, we also make our dataset publicly available. The dataset enables the prediction of diseases based on Bangla symptom inputs, supporting healthcare accessibility for Bengali-speaking populations. Using this dataset, we evaluated multiple machine learning models to predict diseases based on symptoms provided in Bangla and analyzed their performance on our dataset. Both soft and hard voting ensemble approaches combining top-performing models achieved 98\% accuracy, demonstrating superior robustness and generalization. Our work establishes a foundational resource for disease prediction in Bangla, paving the way for future advancements in localized health informatics and diagnostic tools. This contribution aims to enhance equitable access to health information for Bangla-speaking communities, particularly for early disease detection and healthcare interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。