arXiv:2608.10566stat.MLcs.CV2026-08

迭代擦除计数不是概念维度的可靠度量,会受参数变换影响。

Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension

论文配图:Iterative Erasure Count Is Not an Affine-Invariant Concept Dimension
图 1 · 摘自论文原文
  • 提出区分模型定义量与过程定义量,揭示擦除计数依赖测量过程。
  • 在可逆变换下,累积擦除计数从1变2,而预测能力不变。
  • 适用于关注表示几何与测量方法关系的研究者。

神经表示中编码一个概念使用了多少方向?常见做法是反复擦除探测方向,记录停止次数或累计移除秩。我们证明这两个量在信息保持的可逆重参数化下会改变,因此并非内在的概念维度。区分了模型定义的总体量(生成维度、充分线性维度、最小防护秩)与过程定义量(停止次数、累计编辑秩)。在高斯构造中,可逆剪切保持预测问题和三个总体量,但使累计欧氏擦除计数从1变为2。该分离现象对摩尔-彭罗斯最小二乘及所有有限非负岭权重均成立。在两输出全QR流程中,累计编辑秩从2变为环境维数4。相反,当正定度量、探测器、正则化及破缺规则一致传输时,完整累计度量QR轨迹具有仿射等价性;精确协方差仅为推论,非语义度量的规范形式。在已知秩的有限样本Adam/QR校准中,身份混合在全部20次大样本运行中仅接受一次更新,而每个测试剪切a∈{.5,.75,1,1.25,2}均在全部20次运行中接受至少两次更新。冻结V-JEPA2特征的受控重参数化保留零秩预测,却改变实际优化下的后续欧氏轨迹。这些视觉接触实验为压力测试,而非接触维度估计。因此,迭代擦除返回的是由表示几何与完整测量过程共同决定的过程相关估计量,而非独立的语义维度。

原文摘要 · Abstract (English)

How many directions does a neural representation use to encode a concept? A common answer repeatedly erases probe directions and reports the stopping count or cumulative removed rank. We show that both quantities can change under an information-preserving invertible reparameterization, so neither is intrinsically a concept dimension. We distinguish model-defined population quantities (generating dimension, sufficient linear dimension, and minimum guarding rank) from procedure-defined quantities such as stopping count and cumulative edit rank. In a population Gaussian construction, an invertible shear preserves the prediction problem and all three quantities, yet changes the cumulative Euclidean erasure count from one to two. The separation holds for Moore--Penrose ordinary least squares and every finite nonnegative ridge weight. For a two-output full-QR procedure matching our motivating video analysis, cumulative edit rank similarly changes from two to the ambient dimension four. Conversely, the complete cumulative metric-QR trajectory is affine-equivariant when its positive-definite metric, probe, regularizer, and tie-breaking are transported consistently; exact covariance is one corollary, not a canonical semantic metric. In a known-rank finite-sample Adam/QR calibration, identity mixing stops after one accepted update in all 20 large-sample runs, whereas each tested shear $a\in\{.5,.75,1,1.25,2\}$ accepts at least two updates in all 20 runs. Controlled reparameterizations of frozen V-JEPA2 features preserve rank-zero predictions yet alter later Euclidean trajectories under practical optimization. These visual contact experiments are stress tests, not estimates of contact dimension. Iterative erasure therefore returns a procedure-relative estimand jointly determined by representation geometry and the full measurement procedure, not a semantic dimension by itself.

概念分解神经表示可逆变换度量敏感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。