arXiv:2412.19128cs.CVcs.LG2024-12中稿 · ed被引 9

提出语义残差框架,提升多模态统一表征的跨模态泛化能力。

Semantic Residual for Multimodal Unified Discrete Representation

  • 基于语义残差思想设计跨模态信息解耦机制
  • 在跨模态零样本检索上性能显著优于现有模型
  • 适用于需要强跨模态对齐的应用场景

当前多模态统一表征研究主要采用代码本(codebook)形式,通过向量量化(VQ)进行量化,但对其他量化表示形式的探索仍不足。本文提出一种新框架——语义残差跨模态信息解耦(Semantic Residual Cross-modal Information Disentanglement, SRCID),受残差向量量化(RVQ)中数值残差概念启发,利用语义残差实现多模态数据的信息解耦,更好地处理不同模态间的固有差异。该方法增强了统一多模态表征能力,在跨模态泛化与跨模态零样本检索任务上表现优异,平均性能显著超越现有最先进模型,也优于以往基于RVQ和有限标量量化(FSQ)的尝试。

原文摘要 · Abstract (English)

Recent research in the domain of multimodal unified representations predominantly employs codebook as representation forms, utilizing Vector Quantization(VQ) for quantization, yet there has been insufficient exploration of other quantization representation forms. Our work explores more precise quantization methods and introduces a new framework, Semantic Residual Cross-modal Information Disentanglement (SRCID), inspired by the numerical residual concept inherent to Residual Vector Quantization (RVQ). SRCID employs semantic residual-based information disentanglement for multimodal data to better handle the inherent discrepancies between different modalities. Our method enhances the capabilities of unified multimodal representations and demonstrates exceptional performance in cross-modal generalization and cross-modal zero-shot retrieval. Its average results significantly surpass existing state-of-the-art models, as well as previous attempts with RVQ and Finite Scalar Quantization (FSQ) based on these modals.

多模态向量量化跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。