提出可分离记忆与泛化的拓扑测量框架,让神经网络学习机制更清晰。
Measuring Memory and Generalization as Separable Geometric Channels: The Topo^2 Framework
- 用拓扑学方法将记忆与泛化拆分为两个独立通道
- 发现零损失干预可实现泛化上限且几乎不记忆错误标签
- 揭示记忆具有可逆性、可量化、可锚定的数学规律
在噪声标签上训练的深度网络会同时在干净数据上泛化并记住被翻转的标签。传统方法常将二者混为单一容量压力。本文提出 Topo^2 框架,使记忆与泛化在因果上可分离、可测量、受定律支配。表示空间的持久同调 H1 结构分解为类内流形通道(由训练停止点决定)和类间通道(对被记忆翻转样本的单调读出)。干预手段 FM0(从第 0 轮起对翻转样本损失为零)可在达到泛化极限的同时几乎不记忆。在该框架下确立了具分级证据的定律集:(L2) FM0 分离处方(9/9);(L1) 类内通道为训练位置函数(中升趋势 6/6;收敛后 CIFAR-10 3/3,SVHN 2/3);(L3) 环结构恒等式(定义性,非定律);以及 TLS(记忆-泛化拓扑分层):记忆具有因果可加性、锚定性(消除干净样本则表示坍塌)、可逆性(剥离记忆可恢复近极限泛化)和定量可计性(记忆成本定律,有效斜率系数 C ~ 0.38,参考容量下:CIFAR-10 0.3801 / SVHN 0.3806 / CIFAR-100 0.384 / VGG 0.3715,总体依赖容量,并追溯至干净样本特征偏移)。
原文摘要 · Abstract (English)
Deep networks trained on noisy labels simultaneously generalize on clean data and memorize flipped labels. These are usually conflated as pressures on one capacity. We present Topo^2, a measurement framework that makes them causally separable, measurable, and law-governed. Persistent-homology H1 structure of the representation space separates into a within-class manifold channel (a function of the training stopping point) and a cross-class channel (a monotone readout of memorized flipped samples). An intervention, the FM0 prescription (zero loss on flipped samples from epoch 0), reaches each setting's generalization ceiling while memorizing essentially nothing. Within the framework we establish a law set with graded evidence: (L2) FM0 separation prescription (9/9); (L1) the within-channel as a training-position function (mid-rise 6/6; convergence-back CIFAR 3/3, SVHN 2/3); (L3) a ring-construction identity (definitional, not a law); and TLS (memory-generalization topological layering): memory is causally additive, anchored (silencing clean collapses the representation), invertible (stripping memory restores near-ceiling generalization), and quantitatively billable (the memorization cost law, effective slope coefficient C ~ 0.38 at the reference capacity: CIFAR-10 0.3801 / SVHN 0.3806 / CIFAR-100 0.384 / VGG 0.3715, capacity-dependent in general and traced to clean-sample feature displacement). We also publish the framework's boundaries: a falsification ledger of nine dead ends, and an instrument-vindication section that excludes six families of global statistics as explanations of the within-channel. The framework turns "memorization" from an ill-defined capacity into a measurable, separable, invertible topological layer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。