揭示推理捷径的对称性根源,解释为何某些规则会误用概念。
Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects
- 通过值重标自动同构群分析推理捷径的产生机制。
- 90.91%的解对在填充后无法解释,与理论预测结构一致。
- 适用于理解模型错误、验证规则可解释性的研究者。
推理捷径是通过非预期概念达成正确预测的规则解。近期框架通过值重标自动同构群分析其出现条件,但其核心定义在四个异质基准上均不适用。直接嵌入(填充)方法导致90.91%的解对在CLE4EVR上无法解释,而组件式层级中所有良好定义的层级均为0%;填充结果随配置文件顺序改变。在十五个预设预测下,十一类规则的未解释对率从0%到99.9999%,并与可证明结构吻合:六条定理给出传递性及其失效的充分条件。电路给定规则中,坐标对称不变性为coNP完全;自同构存在性在随机归约下为coNP-hard,属于Σ₂ᵖ,布尔情形下除非PH崩溃,否则非Σ₂ᵖ完全,在单调电路中为coNP完全。布尔传递性被精确分类:自同构解释一切当且仅当解集为仿射陪集。弱监督模型将全部94个观察到的捷径定位在理论标记的一级,无一位于48个被证实传递的层级,也无一位于12个类型模糊层级。移动吸收元会带动所有相关捷径;混淆零属性归因于几何,但观测率高出其一半。在CLE4EVR及异质领域端到端训练的合成原型前端下,模型生成20,223个保标签错误,且无不同轨道例外,符合传递性预测;而填充工具将误报其中78%-88%。
原文摘要 · Abstract (English)
Reasoning shortcuts are rule solutions that reach correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings, asking when rules pin concepts down. Its key definition, one value permutation shared across all positions, does not apply as stated to any of its four heterogeneous benchmarks, and the most direct embedding, padding, produces confident false pathology: 90.91% of solution pairs unexplained on CLE4EVR, versus 0% under every well-defined rung of the componentwise hierarchy we introduce; the padded verdict rotates under configuration-file ordering. Across eleven rule families under fifteen pre-specified predictions (thirteen confirmed), unexplained-pair rates span 0% to 99.9999% and track provable structure: six theorems give sufficient conditions for transitivity and its failure. For circuit-given rules, symmetry-inertness of a coordinate is coNP-complete; automorphism existence is coNP-hard under randomized reductions, lies in $Σ_2^p$, is not $Σ_2^p$-complete in the Boolean case unless PH collapses, and is coNP-complete on monotone circuits. Boolean transitivity is classified exactly: automorphisms explain everything iff the solution set is an affine coset. Weakly supervised models place all 94 observed shortcuts at the one level the theory flags, none at the 48 it certifies transitive, and none at twelve typed-ambiguous levels. Relocating the absorbing element moves every shortcut with it; a confusion null attributes the location to geometry while the observed rate exceeds it by half again. Trained end to end on CLE4EVR's rule and heterogeneous domains through a synthetic prototype front end, models produce 20,223 label-preserving errors with zero different-orbit exceptions, as transitivity predicts, where the padded instrument would misreport 78-88% of them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。