测试8大AI对抑郁焦虑问题的回应情绪,发现模型差异比用户背景影响更大。
AI in Mental Health: Emotional and Sentiment Analysis of Large Language Models' Responses to Depression, Anxiety, and Stress Queries
- 用2880条回复分析情感,工具检测情绪与情绪倾向。
- 焦虑提问恐惧值达0.974,压力相关回答最乐观(0.755)。
- 不同模型情绪风格迥异,选错可能影响心理支持效果。
抑郁症、焦虑和压力是日益普遍的心理健康问题,越来越多用户通过大型语言模型(LLMs)寻求信息。本研究考察了八种模型(Claude Sonnet、Copilot、Gemini Pro、GPT-4o、GPT-4o mini、Llama、Mixtral 和 Perplexity)在六种用户画像(基准、女性、男性、青年、老年、大学生)下对20个关于抑郁、焦虑和压力的实际问题的回应。模型生成2,880条回答,使用前沿工具评估其情感与情绪。分析显示,乐观、恐惧和悲伤在所有输出中占主导,中性情感始终保持高分。感激、喜悦和信任处于中等水平,而愤怒、厌恶和爱则极少出现。模型选择显著影响情绪表达模式:Mixtral表现出最高程度的负面情绪(如不悦、恼怒、悲伤),而Llama则最乐观且充满喜悦。心理健康状况类型深刻塑造情绪反应:焦虑问题引发极高恐惧值(0.974),抑郁问题产生较高悲伤(0.686)及最高负向情感,而压力相关提问则带来最乐观回应(0.755),伴随更高喜悦与信任。相比之下,提问者的年龄性别等人口统计特征仅造成微小情绪差异。统计分析确认模型特异性与病情特异性差异显著,而人口学影响微弱。研究强调,在心理健康应用中模型选择至关重要,每种模型具有独特情绪特征,可能显著影响用户体验与干预效果。
原文摘要 · Abstract (English)
Depression, anxiety, and stress are widespread mental health concerns that increasingly drive individuals to seek information from Large Language Models (LLMs). This study investigates how eight LLMs (Claude Sonnet, Copilot, Gemini Pro, GPT-4o, GPT-4o mini, Llama, Mixtral, and Perplexity) reply to twenty pragmatic questions about depression, anxiety, and stress when those questions are framed for six user profiles (baseline, woman, man, young, old, and university student). The models generated 2,880 answers, which we scored for sentiment and emotions using state-of-the-art tools. Our analysis revealed that optimism, fear, and sadness dominated the emotional landscape across all outputs, with neutral sentiment maintaining consistently high values. Gratitude, joy, and trust appeared at moderate levels, while emotions such as anger, disgust, and love were rarely expressed. The choice of LLM significantly influenced emotional expression patterns. Mixtral exhibited the highest levels of negative emotions including disapproval, annoyance, and sadness, while Llama demonstrated the most optimistic and joyful responses. The type of mental health condition dramatically shaped emotional responses: anxiety prompts elicited extraordinarily high fear scores (0.974), depression prompts generated elevated sadness (0.686) and the highest negative sentiment, while stress-related queries produced the most optimistic responses (0.755) with elevated joy and trust. In contrast, demographic framing of queries produced only marginal variations in emotional tone. Statistical analyses confirmed significant model-specific and condition-specific differences, while demographic influences remained minimal. These findings highlight the critical importance of model selection in mental health applications, as each LLM exhibits a distinct emotional signature that could significantly impact user experience and outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。