用可控环境测试大模型做贝叶斯模型选择,发现语义稳定是关键。
Bayesian Wind Tunnels for Model Selection

- 用不变变换构造贝叶斯风洞,实现精确模型选择
- 280万参数模型与最优贝叶斯解熵差仅0.01比特
- 语义稳定性比整数身份更重要,适用于大模型诊断
已有研究证明变压器可在固定假设类中进行精确贝叶斯滤波。能否在数据中识别正确假设类?本文引入模型选择贝叶斯风洞:提供假设类上闭式真后验的受控环境。使用无不动点对合(满足 f(f(x))=x)的280万参数变压器,在整数标记和语义每轮变化的不透明符号下,实现与贝叶斯最优解0.01比特的熵一致性(3次随机种子)。该能力扩展至非嵌套比较:对合与3-循环(互不包含)的类后验平均绝对误差低于0.001,证明超越简单性/子集偏差的真实模型选择。进一步发现显著感知访问条件:当判别统计需算术运算(模加法或乘法),整数标记成功而不透明符号完全失败,且此边界在112倍缩放下(280万→316万参数)仍存在。站位性对照证实:固定重标记的不透明符号可成功(0.009比特MAE),表明稳定语义而非整数身份是电路编译的关键。头部子任务诊断定位失败源于头部反演与算术的组合,而非头部解析本身。对前沿大模型的探测显示定性贝叶斯行为,但校准差距巨大(约55倍),通过有损探测测量,为方向性而非精确结果。
原文摘要 · Abstract (English)
Prior work has shown that transformers can perform exact Bayesian filtering within a fixed hypothesis class. Can they also perform Bayesian model selection -- identifying the correct hypothesis class from data? We introduce model-selection Bayesian wind tunnels: controlled environments where ground-truth posteriors over hypothesis classes are available in closed form. Using fixed-point-free involutions -- whose defining property f(f(x))=x is purely relational -- a 2.8M-parameter transformer achieves 0.01-bit entropy agreement with the Bayesian optimum (3 seeds), with both integer tokens and opaque symbols whose meanings change every episode. This extends to non-nested comparisons: involutions vs. 3-cycles (where neither class is a subset of the other) achieve class-posterior MAE under 0.001, demonstrating genuine model selection beyond simplicity/subset bias. We then identify a sharp perceptual access condition: when the discriminative statistic requires arithmetic -- modular addition (rotations) or multiplication (f(x)=cx mod p) -- model selection succeeds with integer tokens but fails completely with opaque symbols, and this boundary persists under 112x scaling (2.8M to 316M parameters). A stationarity control confirms the operative factor: opaque tokens with a fixed relabeling succeed (0.009-bit MAE), showing that stable semantics, not integer identity, enable circuit compilation. Header subtask diagnostics localize the failure to the composition of header inversion with arithmetic rather than header parsing itself. Probing frontier LLMs on the same tasks shows qualitative Bayesian behavior but a large calibration gap (~55x), measured through lossy probes and therefore directional rather than exact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。