通过微调让大模型模仿抑郁、偏执等病理行为模式,发现其语言生成出现系统性变化。
Modeling Pathology-Like Behavioral Patterns in Language Models Through Behavioral Fine-Tuning

- 用模拟抑郁、偏执行为的数据微调模型,使其在不同情境下稳定选择特定动作。
- 微调后模型在开放任务中对负面和威胁性内容的生成概率显著上升。
- 不同行为模式的模型表现出可区分的响应倾向,适合研究行为与语言的关系。
大型语言模型被广泛用于模拟人类行为。本文提出一种行为诱导框架,通过在结构化决策任务上微调模型,使用受抑郁和偏执等适应不良行为模式启发的合成数据,训练基于Transformer的语言模型在多种情境中持续选择特定类别的行为。随后测试这种行为优化是否引发生成分布的系统性变化。在两种架构下,微调模型均表现出稳定的、跨上下文的语言分布偏移,包括在开放式语言任务中对负面和威胁性解释的概率增加。这些效应超越训练情境,在定性生成结果、心理测量式评估及量化分布指标(如Jensen-Shannon散度)中均可检测到。诱导出的行为特征具有部分特异性:针对不同行为模式优化的模型在评估探针中表现出可分离的反应倾向,表明结构化行为训练产生的是差异化策略级偏差,而非通用分布偏移。研究结果支持将大模型视为基于策略的系统,其中行为约束塑造了涌现表征结构,强调其作为研究行为、解释与生成语言之间关系的可控实验平台的潜力。
原文摘要 · Abstract (English)
Large language models are increasingly used as computational tools for modeling human-like behavior. We introduce a behavioral induction framework that modifies model policies through fine-tuning on structured decision-making tasks: using synthetic datasets inspired by maladaptive behavioral patterns, including depression and paranoia, we train transformer-based language models to consistently select specific classes of actions across diverse contexts. We then test whether this behavioral optimization produces systematic changes in generative distributions. Across two architectures, fine-tuned models show stable, context-general shifts in next-token probability distributions, including increased probability assigned to negative and threat-related interpretations in open-ended language tasks. These effects generalize beyond training contexts and are detectable in qualitative completions, psychometric-style evaluations, and quantitative distributional metrics such as Jensen-Shannon divergence. Induced behavioral profiles also show partial specificity. Models optimized for different behavioral patterns exhibit dissociable response tendencies across evaluation probes, suggesting that structured behavioral training produces differentiated policy-level biases rather than generic distributional skew. We interpret these findings as evidence that consistent behavioral optimization in LLMs can generate stable behavioral and distributional patterns consistent with altered latent priors, linking action selection and language generation. More broadly, the results support a view of LLMs as policy-based systems in which behavioral constraints shape emergent representational structure, highlighting their potential as controlled testbeds for studying the relationship between behavior, interpretation, and generative language in computational models of cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。