arXiv:2606.00906cs.CV2026-06

让小样本医学图像模型用弯曲空间提升性能,效果优于传统方法。

hZACH-ViT: Curved Latent Geometry for Compact Vision Transformers in Low-Data Medical Imaging

论文配图:hZACH-ViT: Curved Latent Geometry for Compact Vision Transformers in Low-Data Medical Imaging
图 1 · 摘自论文原文
  • 用双曲/球面几何替代欧氏空间,优化图像特征表示
  • 在7个数据集上平均提升0.021指标,最大增益达0.055
  • 低曲率配置表现稳定,适合资源受限的医疗场景

紧凑型视觉变换器适用于小样本和资源受限的医学图像任务,但现有方法多假设欧氏潜空间足以组织图像表征。本文提出hZACH-ViT,即ZACH-ViT的曲面几何扩展,该模型去除位置嵌入与类别标记,仅依赖块表示的全局平均池化。为隔离几何作用,保留已验证的ZACH-ViT主干,仅修改最终表示空间与基于原型的分类头,实现欧氏、双曲与球面潜空间的可控对比。在七个MedMNIST数据集上,采用每类50样本、五次随机种子的少样本协议进行评估,共完成770次训练运行,涵盖三类非欧几何、七种曲率值及欧氏基线。所有数据集上最优非欧配置均优于欧氏ZACH-ViT,平均提升0.021(数据集特定主指标),最大提升出现在OCTMNIST(+0.055 MacroF1)。固定低曲率配置在多数数据集上仍保持正向增益,其中曲率c=0.1或0.2的配置占据七项中的六项最优。结果表明,几何与曲率是需依数据集选择的模型变量,且低曲率配置的增益不依赖逐数据集调优。

原文摘要 · Abstract (English)

Compact Vision Transformers are attractive for medical imaging in low-data and resource-constrained settings, but most existing variants assume that Euclidean latent geometry is sufficient for organizing image representations. We introduce hZACH-ViT, a family of curved-geometry extensions of ZACH-ViT, a compact zero-token Vision Transformer that removes positional embeddings and the class token and relies on global average pooling over patch representations. To isolate the role of geometry, we preserve the verified ZACH-ViT backbone and modify only the final representation space and prototype-based classifier head, enabling a controlled comparison between Euclidean, hyperbolic, and spherical latent geometries. We evaluate Poincaré, Klein, and spherical hZACH-ViT heads on seven MedMNIST datasets under an identical few-shot protocol with 50 samples per class and five random seeds. The completed benchmark contains 770 training runs spanning seven datasets, three non-Euclidean geometries, seven curvature magnitudes, and a Euclidean baseline. Across all seven datasets, the best non-Euclidean hZACH-ViT configuration improves over Euclidean ZACH-ViT, with an average gain of +0.021 in the dataset-specific primary metric and the largest improvement on OCTMNIST (+0.055 MacroF1). Fixed low-curvature configurations retain positive gains on the majority of datasets, and low curvature values (c = 0.1 or 0.2) account for six of the seven dataset-level winners. Rather than identifying a universally optimal manifold, our results establish geometry and curvature as dataset-dependent model-selection variables, with fixed low-curvature analyses confirming that gains persist beyond exhaustive per-dataset tuning.

视觉变换器医学图像曲面几何小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。