arXiv:2506.02139cs.AI2025-06

提出统一理论解释大模型如何实现目标行为。

The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning

  • 用锚定强度量化外部信息如何绑定模型内在模式
  • 实验验证跨域锚定、基数变化下的性能阈值现象
  • 提供可优化提示、检索和微调的理论框架

我们提出语义锚定,作为大语言模型将预训练能力转化为目标导向行为的统一解释:外部结构(上下文示例、检索或轻量微调)将模型隐空间中的模式与目标绑定。统一上下文控制理论(UCCT)通过锚定强度 $S = ρ_d - d_r - \ log k$ 形式化该机制,其中 $ρ_d$ 衡量目标在表示空间中的凝聚性,$d_r$ 衡量与先验知识的偏差,$k$ 为锚定预算。UCCT预测性能突变阈值,并严格泛化上下文学习,将检索与微调视为锚定变体。三个受控实验提供证据:实验1展示跨领域锚定可重新绑定强先验,涵盖文本与视觉;实验2在固定复杂度下改变数制(十进制/八进制/九进制),观察到有序阈值与转移模式与 $ρ_d$、$d_r$ 及 $S$ 一致;实验3建立几何与行为关联:层间峰值锚定与轨迹面积可预测少样本阈值 $θ_{50}$。UCCT提供可检验的理论与实用指标,用于优化提示、检索与微调。

原文摘要 · Abstract (English)

We propose semantic anchoring, a unified account of how large language models turn pretrained capacity into goal-directed behavior: external structure (in-context examples, retrieval, or light tuning) binds the model's latent patterns to desired targets. Unified Contextual Control Theory (UCCT) formalizes this via anchoring strength $S = ρ_d - d_r - \log k$, where $ρ_d$ measures target cohesion in representation space, $d_r$ measures mismatch from prior knowledge, and $k$ is the anchor budget. UCCT predicts threshold-like performance flips and strictly generalizes in-context learning, reading retrieval and fine-tuning as anchoring variants. Three controlled studies provide evidence. Experiment 1 demonstrates cross-domain anchoring rebinding strong priors in text and vision. Experiment 2 varies representational familiarity via numeral bases (base-10/8/9) at fixed complexity, yielding ordered thresholds and transfer patterns tracking $ρ_d$, $d_r$, and $S$. Experiment 3 establishes a geometry-to-behavior correlate: layer-wise peak anchoring and trajectory area predict few-shot thresholds $θ_{50}$. UCCT offers testable theory and practical metrics for optimizing prompts, retrieval, and tuning.

大模型认知理论锚定机制推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。