验证对称性先验能否降低样本复杂度,发现错误约束反而有害。
Measuring the Symmetry--Data Exchange Rate

- 设计错误群对照组,排除干扰因素影响
- 对称架构与增强方法在测试时计算对称下性能一致
- 首次量化对称性收益,支持理论预期但需谨慎解读
等变理论预测,架构对称性先验可将样本复杂度降低 |G| 倍;此观点广为引用,但极少在控制条件下以缩放律形式测量。在受控的 C_n 对称任务中,我们报告三项发现:第一,具有相同轨道大小但群类型错误的对照组,在匹配计算量下表现劣于无约束情况(联合成对置信区间 [+0.79, +3.26] 不包含零,多种估计器均稳健),表明错配约束会主动损害性能,而不仅是无效;第二,采用测试时轨道平均的增强基线,与等变模型完全一致——在匹配单元中每轮验证曲线比特级相同,说明架构与增强之间的差距仅在测试时计算不对称时存在,而非绝对;第三,相对交换率 beta_diff = 1.28 在符号和数量级上与理论值 1.0 一致(单层置信区间 [+0.92, +2.05]);但更保守的双层自举法(种子 × 群大小)将其范围扩大至 [-0.63, +1.72],包含零;在 sqrt(2) 间距网格上的细粒度复制实验结果不明确(点估计 -0.82)。方法论贡献——相对速率估计器、错误群对照组及预定义失败分类体系——可推广至任何可参数化的归纳偏置。诚实地指出:主估计量 beta_diff 是事后采纳的,初始分析暴露了正斜率可识别性问题;研究设计未外部预注册;标题数字基于粗网格上七个群大小的 OLS 斜率。本研究为探索性,非确认性测量;最可靠的结果是错误群发现,也是我们最有信心报告的。未来工作包括在新种子下的注册复制。
原文摘要 · Abstract (English)
Equivariance theory predicts that an architectural symmetry prior reduces sample complexity by a factor of |G|; this is widely cited but rarely measured as a scaling law with controls that separate the prior from its confounds. On a controlled C_n-symmetric task, we report three findings. First, a wrong-group control with identical orbit size and matched compute is worse than no constraint (joint pairwise CI [+0.79, +3.26] excludes zero, robust across estimators); misaligned constraint is actively harmful, not merely unhelpful. Second, an augmentation baseline equipped with test-time orbit averaging matches the equivariant model exactly -- bit-identical per-epoch validation curves across matched cells -- so the architecture-vs-augmentation gap is conditional on asymmetric test-time computation, not unconditional. Third, the relative exchange rate beta_diff = 1.28 is consistent in sign and order of magnitude with the theoretical 1.0 (single-level CI [+0.92, +2.05]); the more conservative two-level bootstrap (seeds x group sizes) widens this to [-0.63, +1.72], including zero, and a finer-N replication on a sqrt(2)-spaced grid is inconclusive (point estimate -0.82). The methodological contributions -- the relative-rate estimator that cancels the shared-difficulty confound, the wrong-group control, and a pre-specified failure taxonomy -- transfer to any inductive bias whose strength can be parameterised. Honest scoping: the primary estimator beta_diff was adopted post-hoc after the initial analysis revealed a positive-slope identifiability problem; the design was never externally pre-registered; and the headline number rests on an OLS slope over seven group sizes on a coarse N grid. This is an exploratory study, not a confirmatory measurement; the wrong-group result is the cleanest finding and the one we report with the most confidence. A registered replication on fresh seeds is future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。