让隐空间几何结构直接参与模型优化,提升小样本学习表现。
Geometry as a Missing Axis of Representation Quality: The Variational Geometric Information Bottleneck under Data Scarcity
- 将曲率与内在维度作为显式约束加入瓶颈目标函数
- 在1%~20%标签下,多个数据集上性能优于传统方法
- 特别适合对模型可解释性与泛化能力有要求的研究者
我们研究在数据稀缺情况下,隐空间几何结构作为表示质量的显式组成部分。针对编码器ϕ,定义目标函数Q_{β,γ}(ϕ)=I(ϕ(X);Y)−β𝐶(ϕ)−γd_{int}(ϕ),同时包含任务相关信息、曲率惩罚和内在维度惩罚,使几何特性成为瓶颈准则的一部分,而非事后诊断。在光滑流形、损失迁移和估计器集中假设下,推导出非渐近的低标签泛化界,明确引入内在维度和覆盖复杂度。刻画了信息-几何前沿,并证明了经验代理的一致性。分析揭示编码器几何与学习之间的联系,通过隐空间覆盖数、损失类熵和一致偏差实现。我们将其理论实例化为 exttt{V-GIB},在变分瓶颈训练中加入曲率与维度惩罚。在真实低标签基准测试中,对比了ERM、VIB及多种消融实验,在(1%)至(20%)标签比例下验证其有效性。结果表明,在多个场景下性能提升且几何复杂度降低,尤其在FashionMNIST和CIFAR-10上效果显著,但未发现适用于所有场景的固定正则项。
原文摘要 · Abstract (English)
We study latent geometry as an explicit component of representation quality in data-scarce learning. For an encoder (ϕ), we define (Q_{β,γ}(ϕ)=I(ϕ(X);Y)-β\mathcal C(ϕ)-γd_{\mathrm{int}}(ϕ)), combining task-relevant information with penalties for curvature and intrinsic latent dimension. Thus geometry becomes part of the bottleneck criterion, not only a post hoc diagnostic. Under smooth-manifold, loss-transfer, and estimator-concentration assumptions, we derive non-asymptotic low-label generalization bounds where intrinsic dimension and covering complexity enter explicitly. We characterize the information--geometry frontier and prove empirical-surrogate consistency. The analysis links encoder geometry to learning through latent covering numbers, loss-class entropy, and uniform deviation. We instantiate the theory as \texttt{V-GIB}, adding curvature and dimension penalties to variational bottleneck training. Real low-label benchmarks compare \texttt{V-GIB} with ERM, VIB, and ablations across (1%)--(20%) label fractions. Results show improved performance and reduced geometric complexity in several regimes, especially FashionMNIST and CIFAR-10, while confirming that no fixed regularizer is universally dominant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。