通过因果干预提升大模型内在心智理论能力
CoSToM:Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models

- 用因果追踪定位心智理论关键内部层,揭示其语义编码机制
- 在关键层进行轻量激活调控,显著提升对话中类人社交推理能力
- 适合研究大模型社会智能与内在认知对齐的学者
心智理论(ToM)是社会智能的核心能力,指理解他人心理状态的能力。尽管大语言模型(LLMs)在标准ToM基准上表现良好,但在复杂任务场景中常无法泛化,过度依赖提示工程来模仿推理。这暴露了内部知识与外部行为之间的根本性错位:大模型是否具备内在认知?能否将内部知识稳定外化为高质量行为?为此,我们提出CoSToM(因果导向的心智理论对齐引导框架),从机械解释转向主动干预。首先,采用因果追踪映射内部ToM特征分布,实证发现特定内部层编码了基础的ToM语义。基于此,我们在这些关键层实施轻量级对齐,通过目标激活调控实现优化。实验表明,CoSToM显著提升了类人社交推理能力和下游对话质量。
原文摘要 · Abstract (English)
Theory of Mind (ToM), the ability to attribute mental states to others, is a hallmark of social intelligence. While large language models (LLMs) demonstrate promising performance on standard ToM benchmarks, we observe that they often fail to generalize to complex task-specific scenarios, relying heavily on prompt scaffolding to mimic reasoning. The critical misalignment between the internal knowledge and external behavior raises a fundamental question: Do LLMs truly possess intrinsic cognition, and can they externalize this internal knowledge into stable, high-quality behaviors? To answer this, we introduce CoSToM (Causal-oriented Steering for ToM alignment), a framework that transitions from mechanistic interpretation to active intervention. First, we employ causal tracing to map the internal distribution of ToM features, empirically uncovering the internal layers' characteristics in encoding fundamental ToM semantics. Building on this insight, we implement a lightweight alignment framework via targeted activation steering within these ToM-critical layers. Experiments demonstrate that CoSToM significantly enhances human-like social reasoning capabilities and downstream dialogue quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。