arXiv:2605.08764cs.LGcs.CV2026-05

通过谱稳定化提升小样本医学图像表示学习性能

Anchoring the Eigengap: Cross-Modal Spectral Stabilization for Sample-Efficient Representation Learning

论文配图:Anchoring the Eigengap: Cross-Modal Spectral Stabilization for Sample-Efficient Representation Learning
图 1 · 摘自论文原文
  • 提出有限样本下特征谱的稳定性理论,量化可恢复信号模式数
  • 多模态学习能抑制噪声方向,保持特征值间隔,提升数据效率
  • 适用于小样本医疗影像分析,尤其关注模型可靠性诊断

深度视觉模型在低数据场景下性能急剧下降,尤其是在标注样本稀缺的医学影像中。我们发现这不仅源于过拟合,更根本上是几何失效:有限样本噪声会破坏嵌入协方差,导致特征值间隙塌陷,限制可恢复的信号模式数量。本文建立有限样本表示学习的谱理论,量化了从 N 个样本中可稳定估计的模式数 K(N)。基于扰动理论与浓度界,我们证明仅特征值高于噪声水平 ∥Σ̂ − Σ∥_{op} ∼ √(D/N) 的模式可靠,其对应的截断马氏能量决定分类性能。在幂律谱模型下,该能量可近似为截断黎曼ζ函数,将特征值衰减与数据效率及 AUC 联系起来。在此框架下,多模态学习充当谱稳定器:视觉-语言模型施加低秩约束,抑制噪声主导方向,维持特征值间隙,从而在数据稀缺时提高 K(N)。在 MNIST 与多疾病神经影像数据集上,多模态训练虽未显著提升少样本准确率,但保持了更稳定的特征模式并改善类别分离。结果揭示谱塌陷是低数据学习的根本瓶颈。我们使用截断马氏能量和 K(N) 诊断编码器质量,并提出基于ζ函数的谱滤波作为提升数据效率的合理方法。

原文摘要 · Abstract (English)

Deep vision models degrade sharply in low-data regimes, particularly in medical imaging where labeled samples are scarce. We show this arises not merely from overfitting but from a geometric failure: finite-sample noise corrupts the embedding covariance, collapsing the eigengap and limiting the number of recoverable signal-bearing modes. We develop a spectral theory of finite-sample representation learning that quantifies the recoverable dimension K(N), the number of eigenmodes that can be stably estimated from N samples. Using perturbation theory and concentration bounds, we show that only modes with eigenvalues above the noise floor $\|\hatΣ - Σ\|_{\mathrm{op}} \sim \sqrt{D/N}$ are reliable, yielding a truncated Mahalanobis energy that governs classification performance. Under a power-law spectral model, this energy can be approximated by a truncated Riemann zeta function, linking eigenvalue decay to data efficiency and AUC. Within this framework, multimodal learning acts as spectral stabilization: vision-language models impose low-rank constraints that suppress noise-dominated directions and preserve the eigengap, increasing K(N) under data scarcity. Across MNIST and multi-disease neuroimaging, we show that multimodal training maintains more stable modes and improves class separation, even when unimodal models achieve comparable few-shot accuracy. These results identify spectral collapse as a fundamental bottleneck in low-data learning. We use truncated Mahalanobis energy and K(N) to diagnose encoder quality, and introduce zeta-based spectral filtering as a principled approach to improve data efficiency.

小样本学习多模态谱分析医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。