arXiv:2512.18829cs.AI2025-12

用大模型精准预测精神健康风险等级,提升临床评估准确率。

HARBOR: Holistic Adaptive Risk assessment model for BehaviORal healthcare

  • 构建可感知行为健康的语言模型HARBOR,输出-3到+3的风险分值。
  • 在真实数据上达到69%准确率,远超传统方法与现成大模型。
  • 适用于精神健康长期追踪,适合临床辅助决策系统开发者。

行为健康风险评估因患者数据多模态且情绪状态随时间动态变化而极具挑战。尽管大语言模型(LLMs)展现出强大的推理能力,其在结构化临床风险评分中的效果尚不明确。本文提出HARBOR,一种具备行为健康感知的语言模型,用于预测离散的情绪与风险分数——即港口风险评分(Harbor Risk Score, HRS),取值范围为-3(重度抑郁)至+3(躁狂)。同时发布PEARL数据集,包含三年内三名患者每月的生理、行为及自我报告心理健康信号。我们在多种评估设置和消融实验中对比了传统机器学习模型、商用大模型与HARBOR的表现。结果表明,HARBOR优于经典基线与现成大模型,在测试中达到69%准确率,而逻辑回归仅为54%,最强商用大模型仅29%。

原文摘要 · Abstract (English)

Behavioral healthcare risk assessment remains a challenging problem due to the highly multimodal nature of patient data and the temporal dynamics of mood and affective disorders. While large language models (LLMs) have demonstrated strong reasoning capabilities, their effectiveness in structured clinical risk scoring remains unclear. In this work, we introduce HARBOR, a behavioral health aware language model designed to predict a discrete mood and risk score, termed the Harbor Risk Score (HRS), on an integer scale from -3 (severe depression) to +3 (mania). We also release PEARL, a longitudinal behavioral healthcare dataset spanning four years of monthly observations from three patients, containing physiological, behavioral, and self reported mental health signals. We benchmark traditional machine learning models, proprietary LLMs, and HARBOR across multiple evaluation settings and ablations. Our results show that HARBOR outperforms classical baselines and off the shelf LLMs, achieving 69 percent accuracy compared to 54 percent for logistic regression and 29 percent for the strongest proprietary LLM baseline.

精神健康风险评估大模型应用多模态数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。