arXiv:2604.14593cs.CLcs.AI2026-04

解析大模型如何内嵌复杂情绪,揭示嫉妒的心理机制

Mechanistic Decoding of Cognitive Constructs in Large Language Models

论文配图:Mechanistic Decoding of Cognitive Constructs in Large Language Models
图 1 · 摘自论文原文
  • 用表示工程分离出嫉妒的两个心理根源
  • 8个主流大模型均显示嫉妒呈线性结构组合
  • 可精准检测并抑制有毒情绪,助力AI安全

尽管大语言模型展现出日益复杂的共情能力,其处理复杂情绪的内部机制仍不清晰。现有可解释性方法多将模型视为黑箱或仅关注基础情绪,对更复杂情感状态的认知结构探索不足。为此,我们提出基于表示工程(RepE)的认知逆向工程框架,分析社会比较引发的嫉妒。结合评估理论与子空间正交化、基于回归的加权及双向因果引导,我们分离并量化了嫉妒的两个心理前因:比较对象优越性与领域自我定义相关性,并检验其对模型判断的因果影响。在来自Llama、Qwen和Gemma系列的8个大模型上实验表明,模型原生编码嫉妒为这两个因素的结构化线性组合。其内部表征广泛符合人类心理构念,将优越性视为基础触发因子,相关性作为最终强度放大器。该框架还证明,有毒情绪状态可被机械识别并手术式抑制,为多智能体环境中表征监控与干预提供了可能路径。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex emotions remain unclear. Existing interpretability approaches often treat models as black boxes or focus on coarse-grained basic emotions, leaving the cognitive structure of more complex affective states underexplored. To bridge this gap, we propose a Cognitive Reverse-Engineering framework based on Representation Engineering (RepE) to analyze social-comparison jealousy. By combining appraisal theory with subspace orthogonalization, regression-based weighting, and bidirectional causal steering, we isolate and quantify two psychological antecedents of jealousy, Superiority of Comparison Person and Domain Self-Definitional Relevance, and examine their causal effects on model judgments. Experiments on eight LLMs from the Llama, Qwen, and Gemma families suggest that models natively encode jealousy as a structured linear combination of these constituent factors. Their internal representations are broadly consistent with the human psychological construct, treating Superiority as the foundational trigger and Relevance as the ultimate intensity multiplier. Our framework also demonstrates that toxic emotional states can be mechanically detected and surgically suppressed, suggesting a possible route toward representational monitoring and intervention for AI safety in multi-agent environments.

大模型解释情绪建模认知结构AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。