用语音时间特征提升大模型抑郁检测准确率与可信度
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information
- 引入语音时间特征增强检索生成,替代传统文本检索
- 检测准确率超越纯文本RAG系统,且支持可信度评估
- 无需微调即可达精调模型效果,适合临床辅助场景
大型语言模型(LLMs)在健康领域应用日益广泛,但仅依赖文本输入时在抑郁检测上表现有限。尽管检索增强生成(RAG)通常能提升模型能力,我们实验发现传统文本RAG系统对抑郁检测的提升不显著。这主要源于当前方法未能有效捕捉语音中富含的抑郁相关声学特征。为此,我们系统分析了健康者与抑郁患者的语音时间模式差异,并提出基于语音时间特征的检索增强生成框架SpeechT-RAG。该系统利用语音时间特征实现高精度抑郁检测与可靠的置信度估计,其统一框架在不需额外训练的情况下达到与微调模型相当的性能,同时满足心理健康评估对准确性与可信赖性的双重需求。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been increasingly adopted for health-related tasks, yet their performance in depression detection remains limited when relying solely on text input. While Retrieval-Augmented Generation (RAG) typically enhances LLM capabilities, our experiments indicate that traditional text-based RAG systems struggle to significantly improve depression detection accuracy. This challenge stems partly from the rich depression-relevant information encoded in acoustic speech patterns information that current text-only approaches fail to capture effectively. To address this limitation, we conduct a systematic analysis of temporal speech patterns, comparing healthy individuals with those experiencing depression. Based on our findings, we introduce Speech Timing-based Retrieval-Augmented Generation, SpeechT-RAG, a novel system that leverages speech timing features for both accurate depression detection and reliable confidence estimation. This integrated approach not only outperforms traditional text-based RAG systems in detection accuracy but also enhances uncertainty quantification through a confidence scoring mechanism that naturally extends from the same temporal features. Our unified framework achieves comparable results to fine-tuned LLMs without additional training while simultaneously addressing the fundamental requirements for both accuracy and trustworthiness in mental health assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。