arXiv:2606.15884cs.CL2026-06

分析大模型在法律推理中的关键神经元,发现跨任务共用与特定任务神经元。

Neuron Level Analysis of Large Language Model in Legal Domain Reasoning

论文配图:Neuron Level Analysis of Large Language Model in Legal Domain Reasoning
图 1 · 摘自论文原文
  • 通过神经元归因评分定位并抑制关键神经元
  • 移除共用神经元后仅影响对应任务,验证其任务特异性
  • 法律任务间神经元重叠高,暗示跨司法管辖区的通用特征

我们对七种开源大语言模型在法律领域推理中的神经元层级表现进行了分析,对比了其他应用领域的任务。利用神经元归因分数对神经元进行排序并抑制,发现抑制识别出的关键神经元会显著降低目标任务准确率,而随机抑制相同数量神经元则无此效果。进一步发现,存在一小部分神经元在所有七个任务中均具影响力;一旦这些神经元被移除,后续抑制其余神经元仅导致对应任务性能下降,揭示了每种模型中真正任务特异性的神经元。在法律领域内,三个基准测试表现出较高的神经元重叠度,且往往共同受影响,暗示存在跨越不同司法管辖区的法律相关神经元。实验结果表明,关键神经元是否集中于中间MLP层,可能依赖于输入格式和内容,并非普遍规律。

原文摘要 · Abstract (English)

We presented a neuron-level analysis of legal-domain reasoning in LLMs, comparing it with other applied domain tasks across seven open-weight models. Using neuron attribution scores to rank and suppress influential neurons, we confirmed that suppressing the identified neurons collapses accuracy on the target task, whereas suppressing the same number of random neurons does not. We further found a small subset of neurons influential across all seven tasks; once these are removed, suppressing the remaining neurons degrades only the task they were identified from, revealing genuinely task-specific neurons in every model studied. Within the legal domain, the three benchmarks exhibit relatively high neuron overlap and tend to be affected jointly, suggesting of legal components neurons that span jurisdictions. The distribution of identified neurons in our experiments suggests that the hypothesis that influential neurons are concentrated in middle MLP layers may depend on the input format and content, rather than being a universal phenomenon.

大模型分析法律AI神经元可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。