arXiv:2608.27219cs.CL2026-08中稿 · EMNLP

首个评估大模型长期心理健康感知能力的基准,发现现有模型表现有限。

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

论文配图:BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing
图 1 · 摘自论文原文
  • 构建三大数据集上的长时序心理健康评估基准
  • 零样本模型多数不如简单均值基线,仅强模型或好特征有效
  • 强调需选择性检索历史、扎根时间证据、可解释特征推理

心理状态评估依赖于间断自评量表,将主观感受如压力转化为数值评分,但仅提供稀疏的健康快照。可穿戴设备则能持续、低负担地采集行为与生理信号。近期基于大模型的个人健康代理支持对可穿戴数据的自然语言查询,但多限于短期检索(如一周最高步数)。它们无法评估模型是否能基于长期信号推理并预测心理健康评分,同时给出有依据的理由。为填补此空白,我们提出BALMS——首个系统性的大模型代理在长期心理健康感知中的基准。BALMS涵盖3个真实世界纵向数据集,2类任务(封闭形式的心理健康评分预测与由大模型作为评判者自动评分的推理生成),在5种开源与闭源大模型后端上评估3种代理范式。结果表明,零样本代理极少优于简单均值基线,除非使用更强模型或紧凑、语义明确的特征。思维链提示虽提升推理型模型表现,但无法保证时间定位或数值正确性。结合效率与时间扩展性分析,BALMS凸显出对能选择性检索历史、扎根时间证据、并在可解释行为特征上推理的长期心理健康代理的需求。

原文摘要 · Abstract (English)

Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing. Wearable devices offer longitudinal behavioral and physiological signals for continuous, low-burden monitoring. Recent LLM-driven personal-health agents enable natural language queries over wearable signals, but mainly handle short-term, retrieval-based lookups (e.g., highest step count over a week). They do not evaluate whether agents can reason over long-term signals to predict wellbeing scores paired with evidence-grounded rationales. To address this gap, we introduce BALMS, the first systematic benchmark of LLM-based agentic systems for longitudinal mental health sensing. BALMS spans 3 real-world longitudinal datasets, 2 task families (closed-form wellbeing-score prediction and rationale generation auto-graded by an LLM-as-Judge), 3 agentic paradigms evaluated across 5 open- and closed-source LLM backbones. We find that zero-shot agents rarely outperform a simple mean baseline, except with stronger backbones or compact, semantically meaningful features. Chain-of-thought prompting improves reasoning-oriented backbones, but does not guarantee temporal grounding or numerical correctness. Together with more analysis on efficiency and temporal scaling, BALMS highlights the need for longitudinal mental health agents that selectively retrieve history, ground temporal evidence, and reason over interpretable behavioral features.

大模型代理心理健康长时序建模基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。