发现大模型语义理解具量子式上下文依赖,挑战传统可解释性思路。
The production of meaning in the processing of natural language
- 用CHSH参数检测模型在歧义句中的上下文依赖性
- 模型规模跨度四数量级,上下文依赖性与主流评测无关
- 揭示语义生成本质是动态建构,非静态特征提取
理解自然语言处理中意义生成的基本机制,对设计安全、有思辨性且赋权的人机交互至关重要。若意义是建构而非检索所得,则寻找独立于上下文的特征或神经回路以实现机制可解释性的努力可能从根本上受限。认知科学与社会心理学实验表明,人类语义加工表现出更符合量子逻辑而非经典布尔理论的上下文依赖性;近期研究在大型语言模型中也发现了类似现象——在歧义表达的解释过程中出现贝尔不等式的显著违反。本文探索了跨越四个数量级模型规模的推理参数空间中与该不等式相关的CHSH $|S|$ 参数,并与MMLU、幻觉率及无意义检测基准进行交叉比对。结果发现,$|S|$ 分布的四分位距完全与外部基准正交,整体违反率则与三项基准呈弱负相关。我们研究了$|S|$随采样参数和词序的变化规律,并讨论了真实上下文依赖性对提示注入防御的信息论约束,及其人类类比:大规模构建与维护社会语境依赖性,可在特定解释产生前塑造可能解释的空间。本研究对机制可解释性提出新思考,指出真正的上下文依赖性为语义处理的可分解性设置了信息论上限。
原文摘要 · Abstract (English)
Understanding the fundamental mechanisms governing the production of meaning in the processing of natural language is critical for designing safe, thoughtful, engaging, and empowering human-agent interactions. If meaning is constituted rather than retrieved, then the search for context-independent features or circuits in the pursuit of mechanistic interpretability may be fundamentally limited. Experiments in cognitive science and social psychology have demonstrated that human semantic processing exhibits contextuality more consistent with quantum logical mechanisms than classical Boolean theories, and recent works have found similar results in large language models---in particular, clear violations of the Bell inequality in experiments of contextuality during interpretation of ambiguous expressions. In this work, we explore the CHSH $|S|$ parameter---the metric associated with the inequality---across the inference parameter space of models spanning four orders of magnitude in scale and cross-reference our findings with MMLU, hallucination rate, and nonsense detection benchmarks. We find that the interquartile range of the $|S|$ distribution is completely orthogonal to all external benchmarks, while overall violation rate shows weak anticorrelation with all three benchmarks. We investigate how $|S|$ varies with sampling parameters and word order, and discuss the information-theoretic constraints that genuine contextuality imposes on prompt injection defenses and its human analogue, whereby careful construction and maintenance of social contextuality can be carried out at scale, shaping the space of possible interpretations before any particular one is reached. We consider the implications for mechanistic interpretability and how genuine contextuality sets an information-theoretic bound on the decomposability of semantic processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。