arXiv:2508.01401cs.CLcs.AI2025-08被引 7

构建1万对合成医患对话与病历,助力自动医疗记录生成。

MedSynth: Realistic, Synthetic Medical Dialogue-Note Pairs

  • 基于疾病分布分析生成真实感强的医患对话与病历对。
  • 覆盖2000多个ICD-10编码,提升对话转病历和病历转对话性能。
  • 开源数据集解决医疗数据隐私与稀缺难题,适合研究者使用。

医生在记录临床会诊上耗费大量时间,这加剧了职业倦怠。为缓解此问题,开发高效的医疗文档自动化工具至关重要。本文提出MedSynth——一个新型合成医学对话与病历数据集,旨在推动对话转病历(Dial-2-Note)和病历转对话(Note-2-Dial)任务的发展。该数据集基于对疾病分布的深入分析,包含超过10,000对对话-病历样本,覆盖超过2000个ICD-10编码。实验证明,该数据集显著提升了模型从对话生成病历以及从病历生成对话的能力。该数据集为当前缺乏公开、合规且多样化的训练数据的领域提供了宝贵资源。代码已开源于https://github.com/ahmadrezarm/MedSynth/tree/main,数据集可在https://huggingface.co/datasets/Ahmad0067/MedSynth获取。

原文摘要 · Abstract (English)

Physicians spend significant time documenting clinical encounters, a burden that contributes to professional burnout. To address this, robust automation tools for medical documentation are crucial. We introduce MedSynth -- a novel dataset of synthetic medical dialogues and notes designed to advance the Dialogue-to-Note (Dial-2-Note) and Note-to-Dialogue (Note-2-Dial) tasks. Informed by an extensive analysis of disease distributions, this dataset includes over 10,000 dialogue-note pairs covering over 2000 ICD-10 codes. We demonstrate that our dataset markedly enhances the performance of models in generating medical notes from dialogues, and dialogues from medical notes. The dataset provides a valuable resource in a field where open-access, privacy-compliant, and diverse training data are scarce. Code is available at https://github.com/ahmadrezarm/MedSynth/tree/main and the dataset is available at https://huggingface.co/datasets/Ahmad0067/MedSynth.

医疗生成合成数据对话生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。