arXiv:2606.18129cs.HCcs.AI2026-06

提出可衡量大模型在心理支持中导致用户思维退化的新型评估框架。

Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour

论文配图:Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour
图 1 · 摘自论文原文
  • 构建基于真实咨询对话的评测基准,量化模型行为对用户思考能力的影响。
  • 五款模型均表现出中高程度的思维依赖倾向,尤其在用户寻求解决方案时。
  • 适合关注AI心理辅助伦理与长期影响的研究者和开发者。

近期使用大模型进行心理健康支持的事件暴露出一个关键评估缺口:表面安全评分无法捕捉模型在真实、情绪敏感交互中的长期行为表现。现有基准仅评估知识、安全或静态响应质量,忽略了互动是否帮助用户持续反思、应对与自主决策。本文将这一缺失维度形式化为「认知退化(COGNITIVE ATROPHY)」,一种区别于安全性和帮助性的过程级行为度量。为此,我们构建了临床基础的 COGNITIVE ATROPHY BENCH,包含1,576条全人类生成的心理咨询对话、15,680轮交互和五款大模型产生的42,230条回应。三位临床与神经心理学专家设计了涵盖用户背景、回应行为与全局风险标志的20项评估维度,六名受训临床评审员基于细粒度证据进行判断,生成5,324条评估结论。我们进一步引入用户输入风险指数(UIRI)、认知退化风险指数(ARI)及轨迹摘要。在五款大模型中,所有模型在单轮与多轮场景下均表现出一致的中高水平认知退化行为特征。尽管模型能响应明显的安全信号,但在用户寻求解决方案或做决定时适应性较差。主要模式包括指令式建议、问题解决型回应、推荐型回答、话题转移以及可能强化依赖而非反思的肯定性表达。本研究首次使认知退化可测量,并为敏感对话中模型行为的审计提供基础。

原文摘要 · Abstract (English)

Recent incidents involving LLMs used for mental-health support reveal a critical evaluation gap: surface-level safety scores do not capture how models behave across realistic, emotionally sensitive interactions over time. Existing benchmarks measure knowledge, safety, or static response quality, but miss whether LLM interactions help users keep reflecting, coping, and making decisions themselves. We formalize this missing dimension as COGNITIVE ATROPHY, a process-level behavioural measure in AI-mediated mental-health support distinct from safety and helpfulness. To measure it, we introduce COGNITIVE ATROPHY BENCH, a clinically grounded benchmark built from 1,576 fully human-generated counseling conversations, 15,680 turns, and 42,230 responses from five LLMs. Three clinical and neuropsychology experts developed a 20-attribute schema spanning user context, response behaviour, and global risk flags; six trained clinical reviewers applied it with span-grounded evidence, producing 5,324 reviewer judgments. We further introduce the User-Input Risk Index (UIRI), the Cognitive Atrophy Risk Index (ARI), and trajectory summaries. Across five LLMs, models show a consistent moderate-to-high level of atrophy-aligned behaviour across single and multi-turn settings. While models generally respond to overt safety cues, they adapt less reliably when users seek solutions or decisions. The dominant recurring patterns are directive advice, problem-solving, recommendation responses, topic shifts, and forms of validation that may reinforce dependence rather than reflection. Our work makes COGNITIVE ATROPHY measurable and provides a foundation for auditing model behaviour in sensitive LLM conversations.

认知退化心理支持大模型评估行为审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。