让图文表示真正融合,提升多模态模型性能
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
- 通过多重对齐和训练时融合增强图文对应监督
- 在多个任务上超越主流基线,避免早期饱和
- 训练后移除融合模块,保持推理效率
图像-文本对比预训练已成为视觉表征学习的主流范式,但现有方法生成的表示仍部分保留模态特征。本文提出ITO框架,通过两种协同机制解决此问题:多模态多重对齐通过挖掘多样化的图文对应关系丰富监督信号;轻量级训练时多模态融合模块强制结构化跨模态交互。关键在于该融合模块在推理阶段被移除,保持标准双编码器架构的高效性。大量实验表明,ITO在分类、检索及多模态基准测试中持续优于强基线。分析显示,多重对齐提升判别力,而训练时融合作为关键结构正则化项,消除模态差距并稳定训练动态,防止激进对比学习中的早期饱和现象。
原文摘要 · Abstract (English)
Image-text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality. We propose ITO, a framework addressing this limitation through two synergistic mechanisms. Multimodal multiple alignment enriches supervision by mining diverse image-text correspondences, while a lightweight training-time multimodal fusion module enforces structured cross-modal interaction. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures. Extensive experiments show that ITO consistently outperforms strong baselines across classification, retrieval, and multimodal benchmarks. Our analysis reveals that while multiple alignment drives discriminative power, training-time fusion acts as a critical structural regularizer -- eliminating the modality gap and stabilizing training dynamics to prevent the early saturation often observed in aggressive contrastive learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。