发现大模型会放大数值关联,无关上下文影响其判断
Interpreting Multi-Attribute Confounding through Numerical Attributes in Large Language Models
- 用线性探测与偏相关分析研究模型如何整合多个数值属性
- 模型会系统性放大真实世界的数值关联,且受无关数值干扰
- 揭示了模型在多属性混杂下的决策漏洞,适合关注模型公平性的人
尽管行为研究表明大语言模型存在数值推理错误,但其内在表征机制仍不清晰。我们假设数值属性占据共享的潜在子空间,并探究两个问题:(1) 大模型如何内部整合单个实体的多个数值属性?(2) 无关数值上下文如何干扰这些表征及其下游输出?为此,我们在不同规模的模型上结合线性探测、偏相关分析和基于提示的脆弱性测试。结果表明,大模型编码现实中的数值相关性,但倾向于系统性放大;无关上下文引发幅度表征的一致偏移,对下游输出的影响随模型规模变化。这些发现揭示了大模型决策中的脆弱性,为应对多属性纠缠下的公平、表征感知控制提供了基础。
原文摘要 · Abstract (English)
Although behavioral studies have documented numerical reasoning errors in large language models (LLMs), the underlying representational mechanisms remain unclear. We hypothesize that numerical attributes occupy shared latent subspaces and investigate two questions:(1) How do LLMs internally integrate multiple numerical attributes of a single entity? (2)How does irrelevant numerical context perturb these representations and their downstream outputs? To address these questions, we combine linear probing with partial correlation analysis and prompt-based vulnerability tests across models of varying sizes. Our results show that LLMs encode real-world numerical correlations but tend to systematically amplify them. Moreover, irrelevant context induces consistent shifts in magnitude representations, with downstream effects that vary by model size. These findings reveal a vulnerability in LLM decision-making and lay the groundwork for fairer, representation-aware control under multi-attribute entanglement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。