arXiv:2607.06626cs.LGq-bio.NC2026-07

发现视觉语言模型中与人类快感缺失相关的奖励估值机制。

Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia

论文配图:Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia
图 1 · 摘自论文原文
  • 通过神经科学启发,定位模型中类似伏隔核的奖励预期单元。
  • 扰动这些单元后,模型倾向低努力低回报选择,模拟快感缺失行为。
  • 结果符合临床量表,适用于研究抑郁症机制与AI认知对齐。

近期视觉-语言模型逐渐捕捉到人类认知的复杂方面。本文探究这种对齐是否延伸至奖励估值,基于临床用于评估重度抑郁障碍中快感缺失和动机缺陷的测试,构建了一个机制性分析框架。在大脑中,快感缺失常与伏隔核(NAc)及更广泛的多巴胺奖励系统失调相关。尽管神经影像学已定位这些缺陷,但建立NAc活动与特定行为症状之间的因果联系仍具挑战。我们借鉴神经科学思路,在视觉语言模型中功能性识别出奖励预期单元,并通过针对性扰动检验其因果作用。扰动NAc特异性单元后,模型在努力型决策任务中转向低努力、低回报选项,表现出类似人类快感缺失的行为特征。关键的是,该扰动仅影响奖励估值与预期,而非任务执行能力:当移除基于奖励的选择时,模型表现保持基线水平。诱导出的脆弱性进一步与临床快感缺失量表(如DARS和MAP-SR)一致。综上,本研究揭示了与人类高度相似的奖励估值回路在人工智能模型中的存在。

原文摘要 · Abstract (English)

Recent Vision-Language Models capture increasingly complex aspects of human cognition. Here we ask whether this alignment extends to reward valuation, which we assess in a mechanistic framework built on clinical tests that were developed to evaluate anhedonia and motivational deficits in major depressive disorder. In the brain, anhedonia is frequently linked to dysregulation in the Nucleus Accumbens (NAc) and the broader dopaminergic reward system. While neuroimaging has localized these deficits, establishing a causal link between NAc activity and specific behavioral symptoms remains a challenge. We use these ideas from neuroscience to functionally identify reward-anticipatory units in vision language models, and test their causal role via targeted perturbations. Perturbing NAc-selective units induces behavioral effects that mirror human anhedonia: the model shifts toward low-effort, low-reward options in effort-based decision-making tasks. Crucially, our results reflect a specific deficit in reward valuation and anticipation rather than a loss of task capability: the perturbed model maintains baseline performance when reward-based choice is removed. This induced vulnerability further aligns with clinical anhedonia and motivation scales, including DARS and MAP-SR. Taken together, these results reveal reward valuation circuits in AI models that parallel those in humans.

奖励机制快感缺失视觉语言模型因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。