arXiv:2603.05566cs.LGcs.CL2026-03AAAI被引 1

通过解耦与分布采样,提升图文语义对齐精度。

Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal Alignment

论文配图:Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal Alignment
图 1 · 摘自论文原文
  • 用双路径UNet自适应分离语义与模态特征,加约束确保准确解耦。
  • 提出分布采样方法缓解模态差异,避免对齐偏差与信息丢失。
  • 在多个数据集上超越当前最优方法6.6%至14.2%,适合多模态研究者。

跨模态对齐是多模态学习中的关键任务,旨在实现视觉与语言间的语义一致性,即图像-文本对具有相似语义。传统方法通过追求嵌入一致性来实现语义一致,却忽略了嵌入中包含的非语义信息。直观做法是将嵌入解耦为语义与模态两部分,仅对齐语义部分。但此方法面临两大挑战:(1) 缺乏区分语义与模态信息的明确标准;(2) 模态差距可能导致语义对齐偏差或信息损失。为此,我们提出一种新算法——约束解耦与分布采样(CDDS)。具体包括:(1) 引入双路径UNet,自适应解耦嵌入,并施加多重约束以保证有效分离;(2) 提出分布采样方法,弥合模态差距,确保对齐过程的合理性。在多个基准数据集和模型主干上进行的大量实验表明,CDDS显著优于现有先进方法,性能提升达6.6%至14.2%。

原文摘要 · Abstract (English)

Cross-modal alignment is a crucial task in multimodal learning aimed at achieving semantic consistency between vision and language. This requires that image-text pairs exhibit similar semantics. Traditional algorithms pursue embedding consistency to achieve semantic consistency, ignoring the non-semantic information present in the embedding. An intuitive approach is to decouple the embeddings into semantic and modality components, aligning only the semantic component. However, this introduces two main challenges: (1) There is no established standard for distinguishing semantic and modal information. (2) The modality gap can cause semantic alignment deviation or information loss. To align the true semantics, we propose a novel cross-modal alignment algorithm via \textbf{C}onstrained \textbf{D}ecoupling and \textbf{D}istribution \textbf{S}ampling (CDDS). Specifically, (1) A dual-path UNet is introduced to adaptively decouple the embeddings, applying multiple constraints to ensure effective separation. (2) A distribution sampling method is proposed to bridge the modality gap, ensuring the rationality of the alignment process. Extensive experiments on various benchmarks and model backbones demonstrate the superiority of CDDS, outperforming state-of-the-art methods by 6.6\% to 14.2\%.

跨模态对齐语义解耦分布采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。