通过细粒度物体-短语对齐,提升多模态句子嵌入的准确性。
Improving Multimodal Contrastive Learning of Sentence Embeddings with Object-Phrase Alignment
- 利用分割与目标检测模型提取精准物体-短语对
- 在多种主干模型上,STS任务表现优于强基线
- 适合需要高精度跨模态理解的研究者
多模态句子嵌入模型通常在训练中使用图像-标题对及文本数据。然而,这些配对常包含噪声,如图像或标题中的冗余或无关信息。为缓解此问题,我们提出MCSEO方法,通过引入细粒度物体-短语对齐,增强多模态句子嵌入。具体而言,MCSEO利用现有分割和目标检测模型提取准确的物体-短语对,并基于物体-短语对应关系优化对比学习目标。在不同主干模型上的语义文本相似性(STS)任务实验结果表明,MCSEO始终优于强基线,凸显了精确物体-短语对齐在多模态表征学习中的重要性。
原文摘要 · Abstract (English)
Multimodal sentence embedding models typically leverage image-caption pairs in addition to textual data during training. However, such pairs often contain noise, including redundant or irrelevant information on either the image or caption side. To mitigate this issue, we propose MCSEO, a method that enhances multimodal sentence embeddings by incorporating fine-grained object-phrase alignment alongside traditional image-caption alignment. Specifically, MCSEO utilizes existing segmentation and object detection models to extract accurate object-phrase pairs, which are then used to optimize a contrastive learning objective tailored to object-phrase correspondence. Experimental results on semantic textual similarity (STS) tasks across different backbone models demonstrate that MCSEO consistently outperforms strong baselines, highlighting the significance of precise object-phrase alignment in multimodal representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。