arXiv:2509.13664cs.CLcs.AI2025-09EMNLP被引 5

少数神经元编码问题模糊性,可用来检测并控制大模型回答行为。

Sparse Neurons Carry Strong Signals of Question Ambiguity in LLMs

  • 发现少量神经元(最少一个)在线性编码问题模糊性信号。
  • 基于这些神经元的探测器在跨数据集上表现优于提示与表征基线。
  • 首次实现通过操纵神经元控制模型从回答到回避行为的转变。

现实世界中的问题普遍存在模糊性,但大语言模型(LLMs)常以自信态度作答而未寻求澄清。本文揭示,问题模糊性在线性意义上编码于大模型的内部表示中,并可在神经元层面实现检测与调控。在模型预填充阶段,我们发现仅少数神经元(少至一个)即编码模糊性信息。基于这些模糊性编码神经元(AENs)训练的探测器,在模糊性检测任务中表现优异,且跨数据集泛化能力强,超越基于提示和表征的基线方法。分层分析表明,AENs 源自浅层网络,说明模糊性信号在模型处理流程早期即被编码。最后,我们证明通过操控 AENs 可有效引导模型行为,从直接回答转变为拒绝回答。研究揭示了大模型对问题模糊性的紧凑内生表征,为可解释、可控制的行为提供了新路径。

原文摘要 · Abstract (English)

Ambiguity is pervasive in real-world questions, yet large language models (LLMs) often respond with confident answers rather than seeking clarification. In this work, we show that question ambiguity is linearly encoded in the internal representations of LLMs and can be both detected and controlled at the neuron level. During the model's pre-filling stage, we identify that a small number of neurons, as few as one, encode question ambiguity information. Probes trained on these Ambiguity-Encoding Neurons (AENs) achieve strong performance on ambiguity detection and generalize across datasets, outperforming prompting-based and representation-based baselines. Layerwise analysis reveals that AENs emerge from shallow layers, suggesting early encoding of ambiguity signals in the model's processing pipeline. Finally, we show that through manipulating AENs, we can control LLM's behavior from direct answering to abstention. Our findings reveal that LLMs form compact internal representations of question ambiguity, enabling interpretable and controllable behavior.

大模型模糊性检测神经元分析可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。