arXiv:2608.04234math.STcs.LG2026-08

用联合核熵正则化最优传输实现多模态数据对齐,提升少样本场景下的跨模态检索效果。

Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport

论文配图:Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport
图 1 · 摘自论文原文
  • 通过构建全局相似性核函数,利用细粒度跨模态关系进行结构保持对齐。
  • 在少样本条件下显著提升多模态检索性能,优于现有基线方法。
  • 算法支持大规模计算,可结合已有熵正则最优传输求解器高效运行。

我们研究在缺乏跨模态配对数据情况下,将多模态数据映射到共享表示空间的问题。提出一种结构保持的对齐框架——联合核熵正则化格罗莫夫-沃瑟斯坦最优传输(JK-EGW),通过最小化二次最优传输目标将多个模态映射至共同潜在空间。该方法不依赖原始特征空间距离,而是利用模态内与模态间细粒度相似性关系构建全局亲和性核。理论方面,建立了参数化的样本复杂度率 $n^{-1/2}$,与标准、熵正则及格罗莫夫-沃瑟斯坦最优传输的对应速率一致。算法上,采用可扩展的交替优化过程,通过低秩核近似与变分提升实现熵正则最优传输(EOT)更新,有效缓解二次目标的计算负担,使现有EOT求解器得以应用。实验聚焦于预训练编码器嵌入的后处理对齐,在数据稀缺场景下,所提方法在多模态检索任务中表现优于现有基线。

原文摘要 · Abstract (English)

We study the problem of aligning data from multiple modalities into a shared representation space, focusing on settings where strong pretrained unimodal encoders are available but cross-modal paired data are scarce. We propose a structure-preserving alignment framework, joint kernel entropic Gromov--Wasserstein Optimal Transport (JK-EGW), which maps multiple modalities into a common latent space by minimizing a quadratic optimal transport objective. JK-EGW leverages fine-grained similarity relationships within and across modalities to construct a global affinity kernel instead of relying on raw feature-space distances. Our framework naturally provides explicit control over the geometry and distribution of the latent embedding. On the theory side, we establish parametric sample complexity rate of $n^{-1/2}$, matching the corresponding rates for standard, entropic and Gromov--Wasserstein optimal transport. On the algorithmic side, we derive a scalable alternating procedure to solve JK-EGW with entropic optimal transport (EOT) updates through a low-rank kernel approximation and a variational lifting. This lifting scheme effectively relieves the burden of a quadratic objective, and allowing us to take the advantage of existing EOT solvers. Empirically, we focus on post-hoc alignment of embeddings from pretrained encoders in data-scarce regimes, and show that our proposed method achieves improved multimodal retrieval performance compared to existing alignment baselines.

多模态对齐最优传输少样本学习嵌入对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。