arXiv:2503.10095cs.CLcs.AI2025-03被引 9

用推理方法提升大模型预测心理健康状态的准确性和可解释性

Cognitive-Mental-LLM: Evaluating Reasoning in Large Language Models for Mental Health Prediction via Online Text

  • 引入思维链等推理技术增强大模型对网络文本的分析能力
  • 在多个数据集上显著提升分类性能,最高增益达4.67%
  • 适合关注心理健康的临床研究与可解释人工智能应用者

大型语言模型(LLMs)在从在线文本预测心理健康结果方面展现出潜力,但传统分类方法常缺乏可解释性和鲁棒性。本研究评估了结构化推理技术——思维链(CoT)、自一致性(SC-CoT)和树状思维(ToT),以提升在多个来自Reddit的数据集上的分类准确率。通过零样本和少样本思维链提示策略,使用平衡准确率、F1分数及敏感性/特异性等指标进行分析。结果显示,推理增强的方法相比直接预测有明显提升,尤其在复杂案例中。在Dreaddit数据集上,优于M-LLM 0.52%,优于BERT 0.82%;在SDCNL数据集上,优于M-LLM 4.67%,优于BERT 2.17%。但在抑郁严重程度和CSSRS预测中表现下降,可能源于更广泛的测试集。少样本思维链始终优于其他策略,证实推理驱动模型的有效性。然而,数据集差异凸显模型可靠性和可解释性的挑战。本研究为心理健康文本分类中的推理型大模型提供了全面基准,揭示其在可扩展临床应用中的潜力及未来改进方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated potential in predicting mental health outcomes from online text, yet traditional classification methods often lack interpretability and robustness. This study evaluates structured reasoning techniques-Chain-of-Thought (CoT), Self-Consistency (SC-CoT), and Tree-of-Thought (ToT)-to improve classification accuracy across multiple mental health datasets sourced from Reddit. We analyze reasoning-driven prompting strategies, including Zero-shot CoT and Few-shot CoT, using key performance metrics such as Balanced Accuracy, F1 score, and Sensitivity/Specificity. Our findings indicate that reasoning-enhanced techniques improve classification performance over direct prediction, particularly in complex cases. Compared to baselines such as Zero Shot non-CoT Prompting, and fine-tuned pre-trained transformers such as BERT and Mental-RoBerta, and fine-tuned Open Source LLMs such as Mental Alpaca and Mental-Flan-T5, reasoning-driven LLMs yield notable gains on datasets like Dreaddit (+0.52\% over M-LLM, +0.82\% over BERT) and SDCNL (+4.67\% over M-LLM, +2.17\% over BERT). However, performance declines in Depression Severity, and CSSRS predictions suggest dataset-specific limitations, likely due to our using a more extensive test set. Among prompting strategies, Few-shot CoT consistently outperforms others, reinforcing the effectiveness of reasoning-driven LLMs. Nonetheless, dataset variability highlights challenges in model reliability and interpretability. This study provides a comprehensive benchmark of reasoning-based LLM techniques for mental health text classification. It offers insights into their potential for scalable clinical applications while identifying key challenges for future improvements.

心理健康大模型推理文本分类可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。