用大模型打造可交互的抑郁自测聊天机器人,效果可信且用户更信任。
Development and Evaluation of HopeBot: an LLM-based chatbot for structured and interactive PHQ-9 depression screening
- 用大模型+检索增强生成实现动态提问与实时澄清
- 132人测试显示评分一致性高(ICC=0.91),71%更信任聊天机器人
- 适合需要温和、结构化筛查的群体,尤其关注心理健康服务使用史者
静态工具如患者健康问卷-9(PHQ-9)虽能有效筛查抑郁,但缺乏互动性与适应性。我们开发了HopeBot,一个基于大语言模型(LLM)的聊天机器人,通过检索增强生成与实时澄清功能执行PHQ-9评估。在一项自身对照研究中,132名英中两国成年人完成了自填版与聊天机器人版。评分显示强一致性(ICC = 0.91;45%完全一致)。在75名提供反馈的参与者中,71%表示更信任聊天机器人,因其结构更清晰、解释更易懂、语气更支持性。平均评分(0-10)为:舒适度8.4,语音清晰度7.7,敏感话题处理7.6,建议帮助度7.4;后一项结果在就业状态和既往心理服务使用上差异显著(p < 0.05)。总体87.1%的用户愿重复使用或推荐。结果表明,基于语音的大模型聊天机器人可作为可扩展、低负担的常规抑郁筛查辅助手段。
原文摘要 · Abstract (English)
Static tools like the Patient Health Questionnaire-9 (PHQ-9) effectively screen depression but lack interactivity and adaptability. We developed HopeBot, a chatbot powered by a large language model (LLM) that administers the PHQ-9 using retrieval-augmented generation and real-time clarification. In a within-subject study, 132 adults in the United Kingdom and China completed both self-administered and chatbot versions. Scores demonstrated strong agreement (ICC = 0.91; 45% identical). Among 75 participants providing comparative feedback, 71% reported greater trust in the chatbot, highlighting clearer structure, interpretive guidance, and a supportive tone. Mean ratings (0-10) were 8.4 for comfort, 7.7 for voice clarity, 7.6 for handling sensitive topics, and 7.4 for recommendation helpfulness; the latter varied significantly by employment status and prior mental-health service use (p < 0.05). Overall, 87.1% expressed willingness to reuse or recommend HopeBot. These findings demonstrate voice-based LLM chatbots can feasibly serve as scalable, low-burden adjuncts for routine depression screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。