构建3000场创伤治疗对话合成数据集,助力AI心理援助发展
Thousand Voices of Trauma: A Large-Scale Synthetic Dataset for Modeling Prolonged Exposure Therapy Conversations
- 基于暴露疗法生成6种视角的3000场对话,覆盖500个案例
- 包含20类创伤类型和10种行为表现,真实分布符合临床数据
- 适合研究者开发心理对话模型或训练临床辅助工具
AI心理支持系统的发展受限于治疗对话数据的匮乏,尤其在创伤治疗领域。本文提出「千声创伤」合成数据集,基于创伤后应激障碍(PTSD)的延长暴露疗法,构建了3000场治疗对话。数据包含500个独特案例,每个案例从初始焦虑到峰值痛苦再到情绪处理,呈现6种对话视角。融合年龄18-80岁(均值49.3)、49.4%男性、44.4%女性、6.2%非二元性别,20类创伤类型及10种创伤相关行为,采用确定性与概率性生成方法。分析显示创伤类型分布合理(目击暴力10.6%,欺凌10.2%),症状分布亦具真实性(噩梦23.4%,物质滥用20.8%)。临床专家验证其治疗保真度,认可情感深度并建议优化真实性。同时开发了标准化情感轨迹评估基准。该隐私保护数据集填补了创伤导向心理健康数据空白,为患者应用与临床培训提供宝贵资源。
原文摘要 · Abstract (English)
The advancement of AI systems for mental health support is hindered by limited access to therapeutic conversation data, particularly for trauma treatment. We present Thousand Voices of Trauma, a synthetic benchmark dataset of 3,000 therapy conversations based on Prolonged Exposure therapy protocols for Post-traumatic Stress Disorder (PTSD). The dataset comprises 500 unique cases, each explored through six conversational perspectives that mirror the progression of therapy from initial anxiety to peak distress to emotional processing. We incorporated diverse demographic profiles (ages 18-80, M=49.3, 49.4% male, 44.4% female, 6.2% non-binary), 20 trauma types, and 10 trauma-related behaviors using deterministic and probabilistic generation methods. Analysis reveals realistic distributions of trauma types (witnessing violence 10.6%, bullying 10.2%) and symptoms (nightmares 23.4%, substance abuse 20.8%). Clinical experts validated the dataset's therapeutic fidelity, highlighting its emotional depth while suggesting refinements for greater authenticity. We also developed an emotional trajectory benchmark with standardized metrics for evaluating model responses. This privacy-preserving dataset addresses critical gaps in trauma-focused mental health data, offering a valuable resource for advancing both patient-facing applications and clinician training tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。