构建首个孟加拉语真实远程问诊数据集,助力低资源语言医疗AI发展
DocTalkBN: A Novel Dataset of Expert Telemedicine Conversations in Bengali

- 采集真实远程问诊对话,含音视频与文本对
- 覆盖26个专科,总计170万词元,时长超557小时
- 支持分诊、建议安全评估等临床任务,适合医疗NLP研究
可靠医疗对话AI依赖真实专家-患者互动数据,但此类资源在低资源语言如孟加拉语中仍极为稀缺。本文提出DocTalkBN,一个大规模多模态数据集,来自国家级广播远程问诊节目,包含557.63小时配对音视频与文本、1,515次多轮患者来电、10,274组医生问答,共170万词元,覆盖26个医学专科。与以往基于论坛、书面健康内容或合成数据的资源不同,本数据集保留了低资源环境下真实医疗互动的即兴性、上下文丰富性与口语特征。为支持基准研究,我们从语料中构建三个下游任务:医疗分诊分类、建议安全性评估与医学命名实体识别,并对多种大语言模型与编码器基线进行评测。结果表明,DocTalkBN在临床推理任务中具有实际应用价值。数据集与源码已公开,网址:https://anonymous.4open.science/r/doctalk。
原文摘要 · Abstract (English)
Reliable medical conversational AI requires authentic expert--patient interaction data, yet such datasets remain scarce, especially for low-resource languages such as Bengali. We present DocTalkBN, a large-scale multimodal dataset of real-world expert telemedicine conversations in Bengali, collected from nationally broadcast telemedicine programs featuring board-certified physicians. DocTalkBN contains 557.63 hours of paired audio and text, 1,515 multi-turn patient calls, 10,274 host--doctor question--answer exchanges, totaling 1.7M tokens, spanning 26 medical specialties. Unlike prior resources derived from medical forums, written health content, or synthetic data, our dataset preserves the spontaneity, contextual richness, and spoken characteristics of authentic medical interactions in a low-resource setting. To support benchmark-driven research, we further construct three downstream tasks from the corpus, medical triage classification, advice safety evaluation, and medical named entity recognition, and benchmark a diverse set of large language models and encoder-based baselines. Our results show that DocTalkBN is a practically useful resource, particularly for clinically grounded reasoning tasks. We release this resource to facilitate future research on reliable medical NLP and safer, more culturally grounded healthcare systems for low-resource languages. Our source codes and dataset are publicly available at https://anonymous.4open.science/r/doctalk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。