用低秩微调让大模型自动定位心理治疗关键阶段时间点。
Fine-Tuning Large Audio-Language Models with LoRA for Precise Temporal Localization of Prolonged Exposure Therapy Elements
- 用LoRA微调音视频大模型,处理30秒片段输入。
- 在308个真实治疗会话上误差仅5.3秒,接近人工容忍度。
- 适合临床质控、培训督导,且保护患者隐私。
延长暴露疗法(PE)是治疗创伤后应激障碍的有效方法,但评估治疗师执行一致性需人工审听录音,耗时费力。本文提出一种自动定位核心治疗环节时间边界的方法,直接从音频与转录文本中识别治疗阶段的起止时间。通过低秩适应(LoRA)微调预训练音视频模型Qwen2-Audio,输入聚焦于30秒窗口的音视频-文本对。利用大语言模型生成三类核心阶段标签:治疗师引导(P1)、意象暴露(P2)、暴露后处理(P3),并经专业评分员验证。模型通过任务提示引导的软监督学习,预测归一化的边界偏移。在308个真实治疗会话数据集上,最优配置(LoRA秩为8,30秒窗口)实现平均绝对误差(MAE)5.3秒,处于人工标注可接受范围,具备实际应用价值。进一步分析表明,窗口大小与LoRA秩影响上下文粒度与模型适配效果。本研究构建了一个保护隐私、可扩展的治疗一致性追踪框架,有望支持临床培训、督导与质量控制。
原文摘要 · Abstract (English)
Prolonged Exposure (PE) therapy is an effective treatment for post-traumatic stress disorder (PTSD), but evaluating therapist fidelity remains labor-intensive due to the need for manual review of session recordings. We present a method for the automatic temporal localization of key PE fidelity elements, identifying their start and stop times, directly from session audio and transcripts. Our approach fine-tunes a large pre-trained audio-language model, Qwen2-Audio, using Low-Rank Adaptation (LoRA) to process focused 30-second windows of audio-transcript input. Fidelity labels for three core protocol phases, therapist orientation (P1), imaginal exposure (P2), and post-imaginal processing (P3), are generated via LLM-based prompting and verified by trained raters. The model is trained to predict normalized boundary offsets using soft supervision guided by task-specific prompts. On a dataset of 308 real PE sessions, our best configuration (LoRA rank 8, 30s windows) achieves a mean absolute error (MAE) of 5.3s across tasks, within typical rater tolerance for timestamp review, enabling practical fidelity QC. We further analyze the effects of window size and LoRA rank, highlighting the importance of context granularity and model adaptation. This work introduces a privacy-preserving, scalable framework for fidelity tracking in PE therapy, with potential to support clinician training, supervision, and quality assurance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。