arXiv:2508.08276cs.CLcs.AI2025-08中稿 · the Interplay of M…被引 2

用对比定位法找大模型中关键推理单元,发现效果不如预期。

Evaluating Contrast Localizer for Identifying Causal Units in Social & Mathematical Tasks in Language Models

  • 用对比刺激定位高激活神经单元,再通过删除测试其重要性。
  • 小激活单元有时比高激活单元删了影响更大,打破直觉。
  • 数学和心理理论任务的定位单元会互相干扰,需更精准刺激集。

本研究将神经科学中的对比定位方法应用于大型语言模型(LLMs)和视觉-语言模型(VLMs),旨在识别社会认知(ToM)与数学推理任务中的因果相关神经单元。在11个参数量从3B到90B的LLMs及5个VLM上,通过对比刺激集定位顶部激活单元,并利用靶向消融实验评估其对下游任务准确率的影响。对比功能选择的单元、低激活单元与随机选择单元的删除效果,发现低激活单元有时导致更大的性能下降,且数学任务定位的单元反而更损害心理理论任务表现。结果质疑了基于对比的定位方法的因果有效性,强调需要更广泛、更精准的任务特定刺激集。

原文摘要 · Abstract (English)

This work adapts a neuroscientific contrast localizer to pinpoint causally relevant units for Theory of Mind (ToM) and mathematical reasoning tasks in large language models (LLMs) and vision-language models (VLMs). Across 11 LLMs and 5 VLMs ranging in size from 3B to 90B parameters, we localize top-activated units using contrastive stimulus sets and assess their causal role via targeted ablations. We compare the effect of lesioning functionally selected units against low-activation and randomly selected units on downstream accuracy across established ToM and mathematical benchmarks. Contrary to expectations, low-activation units sometimes produced larger performance drops than the highly activated ones, and units derived from the mathematical localizer often impaired ToM performance more than those from the ToM localizer. These findings call into question the causal relevance of contrast-based localizers and highlight the need for broader stimulus sets and more accurately capture task-specific units.

神经定位大模型解释因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。