arXiv:2604.09645cs.CLcs.AI2026-04中稿 · LREC 2026

用大模型生成高质量荷兰语医疗对话,解决数据隐私难题。

Generating High Quality Synthetic Data for Dutch Medical Conversations

论文配图:Generating High Quality Synthetic Data for Dutch Medical Conversations
图 1 · 摘自论文原文
  • 用微调过的荷兰语大模型,参考真实对话生成合成数据。
  • 生成对话词汇丰富但轮次过于规整,缺乏自然流动感。
  • 适合临床NLP研究者扩展荷兰语医疗数据资源。

医疗对话为临床沟通提供了电子病历中缺失的洞察,但受限于隐私与伦理,领域专用数据集稀缺,制约了可靠临床自然语言处理模型的发展。为此,我们提出一个基于微调荷兰语大模型的合成对话生成流程,以真实医疗对话为语言和结构参考。生成结果通过定量指标和母语者及医疗从业者定性评估。定量分析显示词汇多样性高,但轮次过于规律,呈现脚本化而非自然对话特征。定性评估得分略低于平均,评审指出领域特异性不足和表达不自然问题。量质结果相关性弱,表明仅靠数值指标无法全面反映语言质量。研究证明合成荷兰语医疗对话在技术上可行,但需结合领域知识与精心设计的提示来平衡自然性与结构。本工作为通过伦理方式扩充荷兰语临床NLP资源奠定了基础。

原文摘要 · Abstract (English)

Medical conversations offer insights into clinical communication often absent from Electronic Health Records. However, developing reliable clinical Natural Language Processing (NLP) models is hampered by the scarcity of domain-specific datasets, as clinical data are typically inaccessible due to privacy and ethical constraints. To address these challenges, we present a pipeline for generating synthetic Dutch medical dialogues using a Dutch fine-tuned Large Language Model, with real medical conversations serving as linguistic and structural reference. The generated dialogues were evaluated through quantitative metrics and qualitative review by native speakers and medical practitioners. Quantitative analysis revealed strong lexical variety and overly regular turn-taking, suggesting scripted rather than natural conversation flow. Qualitative review produced slightly below-average scores, with raters noting issues in domain specificity and natural expression. The limited correlation between quantitative and qualitative results highlights that numerical metrics alone cannot fully capture linguistic quality. Our findings demonstrate that generating synthetic Dutch medical dialogues is feasible but requires domain knowledge and carefully structured prompting to balance naturalness and structure in conversation. This work provides a foundation for expanding Dutch clinical NLP resources through ethically generated synthetic data.

医疗对话合成数据大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。