反向映射文本到视觉空间,减少对大规模对齐数据依赖
Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
- 将文本嵌入映射到连续视觉空间,融合于中间Transformer层
- 在九个基准上少监督下表现优异,推理任务提升明显
- 适合追求高效多模态推理的场景,不依赖传统对齐预训练
传统多模态学习依赖对齐预训练,通过大规模图文数据将视觉特征投影至离散文本词元空间。本文重新审视这一设计,提出Inverse-LLaVA:将映射方向反转,将文本嵌入投影至连续视觉表示空间,并在中间Transformer层内完成融合。该表示优先架构无需显式对齐预训练阶段,显著降低对大规模对齐数据的依赖。在九个多模态基准测试中,Inverse-LLaVA在减少监督条件下展现强学习效率,在推理密集型任务上取得显著提升,但在依赖显式视觉-文本对齐的感知任务上出现选择性下降。分析表明,这些权衡主要源于监督范式差异而非架构局限。结果表明,有效的多模态推理并不严格依赖对齐预训练,强调保持连续模态表示的重要性,为解耦表示结构与监督范式提供了新方向,推动更灵活高效的多模态系统设计。
原文摘要 · Abstract (English)
Traditional multimodal learning approaches rely on alignment pre-training to bridge vision and language modalities, typically by projecting visual features into discrete text token spaces using large-scale image--text data. We revisit this design choice and propose Inverse-LLaVA, a multimodal architecture that inverts the conventional mapping direction by projecting text embeddings into continuous visual representation space and performing fusion within intermediate transformer layers. This representation-first design enables effective multimodal reasoning without relying on an explicit alignment pretraining stage and significantly reduces dependence on large alignment datasets. Across nine multimodal benchmarks, Inverse-LLaVA demonstrates strong learning efficiency under reduced supervision, achieving substantial gains on reasoning-intensive tasks while exhibiting selective performance drops on perception tasks that depend on explicit visual--text grounding. Our analysis indicates that these trade-offs primarily reflect differences in supervision regime rather than architectural limitations. Together, these results show that alignment pretraining is not strictly required for effective multimodal reasoning and highlight the importance of preserving continuous modality representations, opening a new direction for multimodal architecture design that decouples representation structure from supervision regime for more flexible and efficient multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。