arXiv:2506.13610cs.CL2025-06被引 7

构建首个孟加拉语疾病症状结构化数据集,助力精准诊断。

A Structured Dataset of Disease-Symptom Associations to Improve Diagnostic Accuracy

  • 从权威医学文献整合疾病与症状的二元关联
  • 覆盖多种疾病,支持机器学习诊断模型训练
  • 填补孟加拉语医疗数据空白,服务弱势语言群体

疾病-症状数据集在医学研究、疾病诊断、临床决策和人工智能健康管理中具有重要意义。本研究系统整理了来自在线资源、医学文献和公开健康数据库的疾病-症状关系,仅纳入同行评审的医学资料,排除非学术来源。数据以表格形式呈现,首列为疾病,其余列为症状,每个症状单元格为二值标记(有/无关联),便于机器学习应用。该结构化数据集可支持疾病预测、临床辅助决策系统及流行病学研究。目前虽已有相关进展,但孟加拉语领域仍缺乏结构化数据。本数据集旨在弥补这一空白,推动多语言医疗信息工具发展,并提升低资源语言人群的疾病预测能力。未来可扩展区域特异性疾病及细化症状关联。

原文摘要 · Abstract (English)

Disease-symptom datasets are significant and in demand for medical research, disease diagnosis, clinical decision-making, and AI-driven health management applications. These datasets help identify symptom patterns associated with specific diseases, thus improving diagnostic accuracy and enabling early detection. The dataset presented in this study systematically compiles disease-symptom relationships from various online sources, medical literature, and publicly available health databases. The data was gathered through analyzing peer-reviewed medical articles, clinical case studies, and disease-symptom association reports. Only the verified medical sources were included in the dataset, while those from non-peer-reviewed and anecdotal sources were excluded. The dataset is structured in a tabular format, where the first column represents diseases, and the remaining columns represent symptoms. Each symptom cell contains a binary value, indicating whether a symptom is associated with a disease. Thereby, this structured representation makes the dataset very useful for a wide range of applications, including machine learning-based disease prediction, clinical decision support systems, and epidemiological studies. Although there are some advancements in the field of disease-symptom datasets, there is a significant gap in structured datasets for the Bangla language. This dataset aims to bridge that gap by facilitating the development of multilingual medical informatics tools and improving disease prediction models for underrepresented linguistic communities. Further developments should include region-specific diseases and further fine-tuning of symptom associations for better diagnostic performance

疾病诊断数据集多语言医疗结构化数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。