arXiv:2504.21800cs.CLcs.AI2025-04EMNLP被引 10

用合成对话模拟创伤治疗,评估其临床真实度。

How Real Are Synthetic Therapy Conversations? Evaluating Fidelity in Prolonged Exposure Dialogues

  • 构建了针对暴露疗法的专用评估指标体系
  • 合成对话在可读性上接近真实(89.2 vs 88.1)
  • 适合临床模型训练者与开发者参考使用

医疗领域中合成数据的采用源于隐私保护、数据获取限制和高标注成本。本文探索以合成的延长暴露(Prolonged Exposure, PE)疗法对话作为创伤后应激障碍(PTSD)临床模型训练的可扩展替代方案。通过语言学、结构和协议特定指标(如轮流说话与治疗忠实度)系统比较真实与合成对话。提出并评估了针对PE的专用指标,提供了一种超越表面流畅性的临床忠实度评估新框架。研究发现,合成数据虽能有效缓解数据稀缺问题并保护隐私,但在捕捉细微治疗动态方面仍具挑战。合成对话成功复现了真实对话的关键语言特征,例如可读性得分相近(89.2 对 88.1),但在关键忠实度指标如痛苦监测方面存在差异。该对比凸显了需采用注重忠实度的评估指标来识别临床显著差异。所提出的模型无关框架为开发者与临床人员在敏感应用部署前评估生成模型忠实度提供了关键工具。研究明确了合成数据可有效补充真实数据的场景,也指出了未来改进方向。

原文摘要 · Abstract (English)

Synthetic data adoption in healthcare is driven by privacy concerns, data access limitations, and high annotation costs. We explore synthetic Prolonged Exposure (PE) therapy conversations for PTSD as a scalable alternative for training clinical models. We systematically compare real and synthetic dialogues using linguistic, structural, and protocol-specific metrics like turn-taking and treatment fidelity. We introduce and evaluate PE-specific metrics, offering a novel framework for assessing clinical fidelity beyond surface fluency. Our findings show that while synthetic data successfully mitigates data scarcity and protects privacy, capturing the most subtle therapeutic dynamics remains a complex challenge. Synthetic dialogues successfully replicate key linguistic features of real conversations, for instance, achieving a similar Readability Score (89.2 vs. 88.1), while showing differences in some key fidelity markers like distress monitoring. This comparison highlights the need for fidelity-aware metrics that go beyond surface fluency to identify clinically significant nuances. Our model-agnostic framework is a critical tool for developers and clinicians to benchmark generative model fidelity before deployment in sensitive applications. Our findings help clarify where synthetic data can effectively complement real-world datasets, while also identifying areas for future refinement.

合成数据心理治疗评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。