arXiv:2504.04238cs.CLcs.AI2025-04

发现极稀疏参数影响大模型心智理论能力,仅扰动0.001%就显著降低推理与理解

Sensitivity Meets Sparsity: The Impact of Extremely Sparse Parameter Patterns on Theory-of-Mind of Large Language Models

  • 定位到影响心智理论的关键稀疏参数,揭示其与位置编码的强关联
  • 扰动0.001%敏感参数即导致心智推理、上下文定位能力大幅下降
  • 适用于关注模型可解释性、社会推理机制与人机交互对齐的研究者

本文从机制角度研究大语言模型(LLMs)中心智理论(ToM)能力的涌现,聚焦极稀疏参数模式的作用。提出新方法识别ToM敏感参数,发现扰动仅0.001%的这些参数便显著削弱ToM表现,并损害上下文定位与语言理解能力。分析表明,这些敏感参数与位置编码模块密切相关,尤其在使用旋转位置编码(RoPE)的模型中,扰动会破坏主导频率激活,影响上下文处理。此外,扰动还通过调节查询与键在位置编码下的夹角,改变注意力机制。该研究深化了对大模型社会推理能力形成机制的理解,连接了人工智能可解释性与认知科学,对提升模型对齐性、缓解偏见及优化人机交互系统具有重要意义。

原文摘要 · Abstract (English)

This paper investigates the emergence of Theory-of-Mind (ToM) capabilities in large language models (LLMs) from a mechanistic perspective, focusing on the role of extremely sparse parameter patterns. We introduce a novel method to identify ToM-sensitive parameters and reveal that perturbing as little as 0.001% of these parameters significantly degrades ToM performance while also impairing contextual localization and language understanding. To understand this effect, we analyze their interaction with core architectural components of LLMs. Our findings demonstrate that these sensitive parameters are closely linked to the positional encoding module, particularly in models using Rotary Position Embedding (RoPE), where perturbations disrupt dominant-frequency activations critical for contextual processing. Furthermore, we show that perturbing ToM-sensitive parameters affects LLM's attention mechanism by modulating the angle between queries and keys under positional encoding. These insights provide a deeper understanding of how LLMs acquire social reasoning abilities, bridging AI interpretability with cognitive science. Our results have implications for enhancing model alignment, mitigating biases, and improving AI systems designed for human interaction.

心智理论可解释性位置编码稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。