探索大模型表示能力在跨模态对齐中的潜力
Representation Potentials of Foundation Models for Multimodal Alignment: A Survey
- 分析大模型在多模态数据中学习到的可迁移表示特性
- 发现视觉、语言、语音等模态表示具结构规律与语义一致性
- 适合关注跨模态对齐与表示学习的研究者阅读
大模型通过大规模异构数据预训练,习得高度可迁移的表示。越来越多研究指出,这些表示在不同架构和模态间表现出显著相似性。本文调研大模型的表示潜力,即其学习表示在单模态中捕捉任务特异性信息的能力,以及在多模态间实现对齐与统一的可迁移基础。我们回顾代表性大模型及衡量对齐的关键指标,综合来自视觉、语言、语音、多模态及神经科学领域的实证证据。结果表明,大模型的表示空间常呈现结构规律与语义一致性,使其成为跨模态迁移与对齐的有力候选。进一步分析促进表示潜力的关键因素,讨论开放问题并指出潜在挑战。
原文摘要 · Abstract (English)
Foundation models learn highly transferable representations through large-scale pretraining on diverse data. An increasing body of research indicates that these representations exhibit a remarkable degree of similarity across architectures and modalities. In this survey, we investigate the representation potentials of foundation models, defined as the latent capacity of their learned representations to capture task-specific information within a single modality while also providing a transferable basis for alignment and unification across modalities. We begin by reviewing representative foundation models and the key metrics that make alignment measurable. We then synthesize empirical evidence of representation potentials from studies in vision, language, speech, multimodality, and neuroscience. The evidence suggests that foundation models often exhibit structural regularities and semantic consistencies in their representation spaces, positioning them as strong candidates for cross-modal transfer and alignment. We further analyze the key factors that foster representation potentials, discuss open questions, and highlight potential challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。