arXiv:2506.21096cs.CL2025-06ACL被引 5

通过双层对齐提升图文句子表示质量,解决跨模态偏差与模态内语义分歧问题。

DALR: Dual-level Alignment Learning for Multimodal Sentence Representation Learning

  • 设计一致性学习模块,细化图文对齐,缓解负样本干扰。
  • 引入排序蒸馏与全局模态内对齐,捕捉句子间复杂关系。
  • 在STS和迁移任务中优于当前最优方法,适合多模态语义理解场景。

现有多模态句子表示学习方法虽表现优异,但多数仅关注图像与文本的粗粒度对齐,面临跨模态错位偏差和模态内语义发散两大挑战,显著降低句子表示质量。为此,本文提出DALR(Dual-level Alignment Learning for Multimodal Sentence Representation)。针对跨模态对齐,设计一致性学习模块,软化负样本并利用辅助任务的语义相似性实现细粒度对齐;同时认为句子关系超越二元正负标签,具有更复杂的排序结构,因此融合排序蒸馏与全局模态内对齐学习以增强表示能力。在语义文本相似性(STS)和迁移(TR)任务上的全面实验验证了该方法的有效性,持续优于当前最优基线。

原文摘要 · Abstract (English)

Previous multimodal sentence representation learning methods have achieved impressive performance. However, most approaches focus on aligning images and text at a coarse level, facing two critical challenges:cross-modal misalignment bias and intra-modal semantic divergence, which significantly degrade sentence representation quality. To address these challenges, we propose DALR (Dual-level Alignment Learning for Multimodal Sentence Representation). For cross-modal alignment, we propose a consistency learning module that softens negative samples and utilizes semantic similarity from an auxiliary task to achieve fine-grained cross-modal alignment. Additionally, we contend that sentence relationships go beyond binary positive-negative labels, exhibiting a more intricate ranking structure. To better capture these relationships and enhance representation quality, we integrate ranking distillation with global intra-modal alignment learning. Comprehensive experiments on semantic textual similarity (STS) and transfer (TR) tasks validate the effectiveness of our approach, consistently demonstrating its superiority over state-of-the-art baselines.

多模态句子表示对齐学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。