用能量模型提升图文对齐,让跨模态融合更精准。
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems
- 基于能量模型构建图文联合嵌入架构,自监督学习更灵活。
- 在多任务评测中达到顶尖性能,尤其在情感分析上领先。
- 适合做视觉问答等跨模态应用,推动下游任务发展。
本文聚焦人工智能中的多模态对齐问题,尤其是文本与图像模态间的语义鸿沟影响了多模态融合效果。为此,提出文本-图像联合嵌入预测架构(TI-JEPA),一种基于能量模型(EBM)框架的创新预训练策略,以捕捉复杂的跨模态关系。TI-JEPA利用能量模型在自监督学习中的灵活性,增强文本与视觉元素的兼容性。在多个基准测试上的大量实验表明,TI-JEPA在多模态情感分析任务上达到当前最优表现,并有望适用于广泛多模态任务(如视觉问答)。研究结果表明,能量模型框架在推进多模态融合方面具有巨大潜力,可显著提升下游应用效果。
原文摘要 · Abstract (English)
This paper focuses on multimodal alignment within the realm of Artificial Intelligence, particularly in text and image modalities. The semantic gap between the textual and visual modality poses a discrepancy problem towards the effectiveness of multi-modalities fusion. Therefore, we introduce Text-Image Joint Embedding Predictive Architecture (TI-JEPA), an innovative pre-training strategy that leverages energy-based model (EBM) framework to capture complex cross-modal relationships. TI-JEPA combines the flexibility of EBM in self-supervised learning to facilitate the compatibility between textual and visual elements. Through extensive experiments across multiple benchmarks, we demonstrate that TI-JEPA achieves state-of-the-art performance on multimodal sentiment analysis task (and potentially on a wide range of multimodal-based tasks, such as Visual Question Answering), outperforming existing pre-training methodologies. Our findings highlight the potential of using energy-based framework in advancing multimodal fusion and suggest significant improvements for downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。