用语音+文本评估患者焦虑程度,发现声音信息能补足文字遗漏的关键线索。
CBT-Audio: Evaluating Audio Language Models for Patient-Side Distress Intensity Estimation in CBT Session Recordings

- 构建语音与文本双模态数据集,支持心理治疗中的情绪评估
- 8/10模型在加入语音后表现提升,尤其在言辞与语气不一致时效果显著
- 适合研究心理健康AI、语音情感分析及人机交互的学者和开发者
认知行为疗法常通过口语对话进行,治疗师不仅关注患者说什么,还重视其表达方式,这些非语言线索对判断患者心理状态至关重要。现有研究多基于文本,因公开可用的语音数据受限于伦理与隐私问题。为此,我们提出CBT-Audio数据集,包含96段公开可获取的CBT录音中1,802个患者发言片段,每段配有专家标注的逐轮痛苦程度标签。我们评估了10个开源语音语言模型在三种输入条件下的表现:仅音频、仅文本,或音频+文本。结果表明,音频可提供文本之外的有效信息,尤其在与文本结合时:8个模型家族在加入音频后性能优于仅使用文本,其中4个提升显著;案例分析显示,当言语内容与声调不一致时,语音带来的增益最明显。CBT-Audio使患者语音行为可量化评估,为心理医疗中音频语言模型的发展提供支持。
原文摘要 · Abstract (English)
Cognitive behavioural therapy is widely used to help patients understand and manage psychological distress. It is often delivered through spoken conversation, where therapists attend not only to what patients say, but also to how they say it, because these cues can help therapists decide how to respond and adapt treatment. Progress in building AI systems for CBT remains largely limited to text, partly because most available datasets are text based and shareable spoken CBT data are scarce under ethical and privacy constraints. This creates a blind spot because text based models and evaluations cannot capture the mismatch between the transcript and the patient's voice, even though therapists often rely on this mismatch to understand patient distress. We introduce CBT-Audio, a dataset for evaluating patient distress estimation from spoken CBT sessions with audio language models. CBT-Audio contains 1,802 patient turns from 96 publicly available CBT recordings, with turn-level distress labels validated on an experts-annotated subset. We evaluate 10 open source audio language models under three input conditions, where models receive only patient audio, only the transcript, or both audio and transcript. Our results show that audio can provide useful information beyond text, especially when combined with transcripts. Adding audio to transcript input improves distress estimation over using the transcript alone in 8 of 10 model families, with significant gains in 4, and case studies show the clearest benefit when verbal content and vocal delivery diverge. CBT-Audio makes spoken patient behaviour measurable for AI evaluation in CBT-related tasks and supports future work on audio language models for mental health interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。