arXiv:2604.26370cs.CVcs.LG2026-04

用拓扑结构提升视觉语言模型在小样本下的泛化能力

Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning

论文配图:Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning
图 1 · 摘自论文原文
  • 基于持久同调识别跨模态关键连接边并进行对齐
  • 在遥感与时尚检索上分别取得显著与稳定提升
  • 适合关注多模态结构建模与小样本学习的研究者

视觉语言模型虽性能优异,但在专业领域泛化能力差。半监督视觉语言学习通过少量标注图像-文本对和大量未标注图像缓解此问题,但现有方法多为成对匹配,难以捕捉多模态表示流形的全局结构。现有基于拓扑的方法依赖持久图匹配,既无法保证几何对齐,也未利用图像-文本配对信息。本文提出拓扑感知多模态表示对齐(ToMA),利用持久同调识别拓扑显著边,并通过已有跨模态对应关系对齐不同模态。ToMA同时利用H_0-death边与轻量级H_1-birth边,无需构建2-单形即可捕获连通性与环状结构。实验表明,ToMA在遥感任务上实现显著提升,在时尚检索上取得稳健增益;分析显示其优于其他拓扑目标,且轻量级H_1-birth边提供了有效的高阶结构信号。

原文摘要 · Abstract (English)

Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language learning mitigates this limitation by leveraging a small set of labeled image-text pairs together with abundant unlabeled images, existing methods remain fundamentally pairwise and fail to model the global structure of multimodal representation manifolds. Existing topology-based alignment methods rely on persistence diagram matching, which neither guarantees geometric alignment nor utilizes the image-text pairing information central to vision-language learning. We propose Topology-Aware Multimodal Representation Alignment (ToMA), a framework that uses persistent homology to identify topologically salient edges and aligns them across modalities through available cross-modal correspondences. ToMA leverages both H_0-death edges and lightweight H_1-birth edges, allowing it to capture both connectivity and cycle structure without constructing 2-simplices. Experiments show that ToMA yields stable gains, with clear improvements on remote sensing and modest but consistent benefits on fashion retrieval. Additional analysis shows that ToMA is more stable than alternative topology-based objectives and that lightweight H_1-birth edges provide useful higher-order structural signals.

多模态学习拓扑分析半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。