arXiv:2607.18691cs.AIcs.CL2026-07

用自然语义元解释大模型情绪,效果优于传统方法。

Semantic Primes as Explanans for Emotion in Large Language Models

论文配图:Semantic Primes as Explanans for Emotion in Large Language Models
图 1 · 摘自论文原文
  • 以自然语义元作为模型内部情绪解释变量。
  • 干预语义元可使情绪控制强度提升三倍,选择性翻倍。
  • 语义元解释可替代情绪表达,更具科学合理性。

大语言模型(LLMs)的情绪机制研究已有进展,但如何解释其情绪表现仍不明确。尽管情绪表征、组件和电路可被恢复,但这些解释常陷入循环论证,且维度任意、无终止性。本文提出使用自然语义金属语言(NSM)的语义元作为更基础的解释变量。在四个指令微调的LLM(Llama-1B、Gemma-2B、Gemma-9B、OLMo-7B)上实验发现:(1)NSM语义元是可恢复的内部元素;(2)在参考模型中,基于语义元方向的干预比最优评估方向的情绪控制强度高约三倍,选择性高两倍;(3)模型将语义元解释视为与对应情绪表达等价。这些证据表明,根据科学解释标准,NSM语义元比诸多替代方案更适合作为LLM情绪的解释项。

原文摘要 · Abstract (English)

Progresses have been made on understanding emotion mechanisms of large language models (LLMs). However, how to explain emotion in LLMs, or even what constitutes good explanations, are less clear. Emotion representations, components, circuits are widely recoverable, but as explanations of a model's own computation they are circular; the emotion space dimensions tend to be arbitrary and non-terminating. A pressing question to ask is whether a more primitive set of internal variables does the work: the semantic primes of the Natural Semantic Metalanguage (NSM). Across four instruction-tuned LLMs (Llama-1B, Gemma-2B, Gemma-9B, OLMo-7B), experiments show that the NSM primes are (1) recoverable internal elements; and (2) on the reference model, intervening with a prime based direction controls emotion about three times as strongly, and twice as selectively, as the best appraisal based direction; and (3) the model treats a prime based explication as interchangeable with the corresponding emotion. These evidences suggest that NSM primes seem to be better explanans for emotion in LLMs than many alternative options according to scientific explanations criteria.

情绪解释语义元大模型机制解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。