arXiv:2509.24431cs.LG2025-09

通过模态对齐实现多模态嵌入的语义压缩,大幅节省存储且不损失性能。

Semantic Compression via Multimodal Representation Learning

  • 利用模态对齐使跨模态嵌入共享语义空间,以质心替代多嵌入。
  • 在多个大规模下游任务中实现显著压缩率,性能保持不变。
  • 适合需要高效部署多模态模型的研究者与工程团队。

多模态表示学习生成高维嵌入,在共享潜在空间中对齐不同模态。尽管这增强了泛化能力,但也带来存储和下游处理的可扩展性挑战。核心开放问题是如何实现语义压缩,即在保留跨模态共享语义内容的前提下减少多模态嵌入的内存占用。本文证明了缩小模态差距(即不同模态嵌入间的残余分离)与训练后语义压缩可行性之间存在强关联。当差距足够小时,表达相同语义的跨模态嵌入会共享空间中的共同部分,其质心可作为该语义概念的忠实表示。因此可用单一质心替换多个嵌入,实现显著内存节省。我们提出一种基于此直觉的新方法,直接作用于预训练编码器。在多种大规模多模态下游任务中验证了其有效性。结果表明,模态对齐是语义压缩的关键前提,所提方法在不牺牲性能的情况下实现显著压缩。

原文摘要 · Abstract (English)

Multimodal representation learning produces high-dimensional embeddings that align diverse modalities in a shared latent space. While this enables strong generalization, it also introduces scalability challenges, both in terms of storage and downstream processing. A key open problem is how to achieve semantic compression, reducing the memory footprint of multimodal embeddings while preserving their ability to represent shared semantic content across modalities. In this paper, we prove a strong connection between reducing the modality gap, which is the residual separation of embeddings from different modalities, and the feasibility of post-training semantic compression. When the gap is sufficiently reduced, embeddings from different modalities but expressing the same semantics share a common portion of the space. Therefore, their centroid is a faithful representation of such a semantic concept. This enables replacing multiple embeddings with a single centroid, yielding significant memory savings. We propose a novel approach for semantic compression grounded on the latter intuition, operating directly on pretrained encoders. We demonstrate its effectiveness across diverse large-scale multimodal downstream tasks. Our results highlight that modality alignment is a key enabler for semantic compression, showing that the proposed approach achieves significant compression without sacrificing performance.

多模态嵌入压缩语义对齐模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。