arXiv:2605.29900cs.LGcs.IT2026-05

提出一种新框架,让多模态数据自动对齐并保留关键信息。

OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment

论文配图:OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment
图 1 · 摘自论文原文
  • 以信息瓶颈原理为基础,用‘一 vs 其余’视角建模多模态关系。
  • 在多个基准上表现优异,尤其在跨模态检索和无模态偏好任务中领先。
  • 无需调参的几何感知投影,适合任意数量模态的对齐场景。

对比学习在配对模态对齐中有效,但多于两模态的对齐仍具挑战且研究较少。传统成对对比损失将多模态对齐分解为独立二元比较,无法显式建模多模态间的高阶依赖。近期方法从统计或几何角度尝试解决,但任意模态对齐仍缺乏明确准则来定义各模态应保留与压缩的信息。本文基于信息瓶颈原则重新审视任意模态对齐:充分性要求保留可由其余模态预测的信息,最小性要求压缩未被其余模态支持的模态特有信息。这自然导向‘一 vs 其余’视角。我们提出 OVA-IB,一种面向任意模态对齐的信息瓶颈框架。OVA-IB 优化一个可计算的‘一 vs 其余’对比下界以实现充分性,结合类总相关性的双重目标;使用无需参数的几何感知投影评分;并通过上界正则化项,限制每个表示对其自身输入的依赖,该依赖由其余模态诱导的表示分布所界定。在分类、回归、模态无关评估及跨模态检索等基准上的实验表明,该方法具有强而稳健的性能。

原文摘要 · Abstract (English)

Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively underexplored. Pairwise CLIP-style losses decompose multi-modal alignment into independent two-way comparisons and therefore do not explicitly model higher-order dependencies among multiple modalities. Recent beyond-pairwise objectives approach this problem from statistical or geometric perspectives, but arbitrary-modality alignment still lacks a principled criterion for defining what each modality should preserve and compress relative to the others. We revisit arbitrary-modality alignment through the Information Bottleneck principle. In multi-modal learning, sufficiency should preserve information predictable from the remaining modalities, while minimality should compress modality-specific information not supported by them. This naturally leads to a One-vs-All view, where each modality is characterized with respect to the remaining modalities. We propose OVA-IB, an Information Bottleneck framework for arbitrary-modality alignment. OVA-IB optimizes a tractable One-vs-All contrastive lower bound for sufficiency connected to a Dual Total Correlation-style objective, uses a parameter-free geometry-aware projection score, and derives a tractable upper-bound regularizer for minimality by bounding each representation's dependence on its own input with representation distributions induced by the remaining modalities. Experiments on classification, regression, modality-agnostic evaluation, and cross-modal retrieval benchmarks demonstrate strong and robust performance.

多模态信息瓶颈对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。