arXiv:2606.30491cs.CLcs.AI2026-06

构建可控制的医患对话模拟框架,用于训练和评估医疗沟通分析系统。

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

  • 基于预设场景与行为代码本生成可控医患对话数据。
  • 生成3388段对话,语音自然度与转录准确率均达可用水平。
  • 适合开发医疗沟通分析模型的研究者及系统验证团队使用。

环境数字记录器的广泛应用推动了大量医患对话数据的采集。人工标注临床沟通数据成本高、不一致且难以扩展,促使采用AI驱动的沟通标注系统。然而,评估这些系统需真实对话与人工标注标签,二者均难大规模获取。为此,我们开发了SIMAX(多保真度与标注医患对话模拟的可扩展可解释框架),可从预设临床场景、人物角色、语音条件和目标沟通行为生成受控的医患对话,并通过全局代码本(Global Codebook)和具体行为代码本(WISER Codebook)实现行为控制。在三个专科领域生成3,388段对话,涵盖多个就诊阶段、人物特征与口音条件。自动评估显示平均UTMOS为3.03,WV-MOS为2.61,词错误率(WER)和字符错误率(CER)分别为0.07和0.05,CLAP余弦相似度为0.41,表明语音自然度合理、转录准确率高、文本与音频对应良好。人类评估中,平均感知质量评分(MOS)为4.67,临床现实感中位评分为3.00。下游评估表明,SIMAX可检测通信编码系统对行为目标的响应能力,并揭示部分维度敏感性不足的问题。研究结论:SIMAX可生成可控、可复现的模拟医患对话,为开发、验证与优化沟通编码系统提供数据基础。

原文摘要 · Abstract (English)

Background. The widespread deployment of ambient digital scribes is driving large-scale capture of clinician-patient dialogues. Human coding of clinical communication data remains costly, inconsistent, and difficult to scale, motivating AI-driven communication coding systems. However, evaluating these systems requires real-world dialogues and human-coded labels, both hard to obtain at scale. Methods. We developed SIMAX (Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation), a framework for generating controlled clinical dialogue data with reference behavioral annotations. SIMAX generates clinician-patient dialogues from predefined clinical scenarios, personas and voice conditions, and target communication behaviors. Behaviors are controlled using two codebooks: the Global Codebook for overall communication quality and the WISER Codebook for specific countable behaviors. We evaluated SIMAX using automated and human quality assessments and an example communication coding system. Results. SIMAX generated 3,388 simulated dialogues across three specialties, multiple visit stages, persona characteristics, and accent conditions. Automated assessment showed mean UTMOS and WV-MOS scores of 3.03 and 2.61, WER and CER of 0.07 and 0.05, and CLAP cosine similarity of 0.41, suggesting reasonable speech naturalness, high transcription fidelity, and positive text-audio correspondence. Human evaluation showed a median MOS of 4.67 and a median clinical realism score of 3.00. Downstream evaluation suggests that SIMAX can assess how a communication coding system responds to behavioral targets and reveal insufficient sensitivity in some dimensions. Conclusions. SIMAX generates controlled and reproducible simulated clinician-patient dialogues, providing a data foundation for developing, validating, and refining communication coding systems.

对话模拟医疗沟通可解释性多保真度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。