arXiv:2608.27651cs.LG2026-08

用几何对称性诊断模型设计缺陷,避免数据冗余却无法识别的问题。

More Data Cannot Break a Symmetry: Identifiability by Design

  • 基于刺激几何的对称性设计诊断,提前发现不可辨识的结构缺陷。
  • 相同数据量下,对称设计失败率75%,非对称设计成功率98%。
  • 仅凭几何诊断选颜色,无需学习即可避免灾难性对齐失败。

无监督表征对齐仅靠几何恢复刺激对应关系,但刺激几何的自同构群限制了任何对齐能识别的内容。已知的对称性问题(Demetci等,2024)在密集采样时因近似重复导致重标记代价极低,传统诊断方法会误判设计优劣。本文将此不变性转化为设计阶段的诊断工具与干预手段。在颜色空间中,候选几何具解析形式,揭示其结构性失败:64倍重启预算下对称设计完全不变,而相同规模的非对称集总能成功恢复。表征模型判别与对应恢复基本无关(r = -0.02,3,000子集)。仅依此诊断选择9种颜色,不依赖任何学习表征,即可使93个模型表征远离退化点,将灾难性对齐失败从75%降至2%,模型、层、样本数和求解器均固定。该风险在均匀分布方向、音调或运动方向等场景同样存在,且检测成本仅为一次函数调用,可在数据收集前完成。

原文摘要 · Abstract (English)

Unsupervised representational alignment recovers a stimulus-by-stimulus correspondence from geometry alone, but the automorphism group of the stimulus geometry bounds what any such alignment can identify, before data exist. The obvious diagnostic for this degeneracy, the cheapest non-identity relabelling, ranks two published designs in the wrong order, because dense sampling creates near-duplicates whose transposition is nearly free. We turn this known invariance (Demetci et al., 2024) into a design-time diagnostic and intervention. In colour, where candidate geometries have closed form, we show that the failure is structural: sixty-four times the restart budget leaves a symmetric design unmoved while an asymmetric set at the same N recovers every time. Discriminating representational models and recovering a correspondence are essentially uncorrelated objectives (r = -0.02 over 3,000 subsets). Choosing nine colours by this diagnostic alone, without consulting any learned representation, moves all 93 model representations away from the degenerate point and cuts catastrophic alignment failures from 75% to 2% with the models, the layers, N and the solver all held fixed. The same risk arises wherever a regular design meets its candidate geometry's isometry group, including evenly spaced orientations, tones, or motion directions, and the check costs one function call before data collection.

表征对齐几何对称性模型设计可辨识性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。