用多头注意力增强图文匹配,提升跨模态理解能力。
Multi-Head Attention Driven Dynamic Visual-Semantic Embedding for Enhanced Image-Text Matching
- 引入多头自注意力机制并行捕捉视觉与语义的多重关系。
- 在Flickr30k数据集上,图文检索准确率优于现有方法。
- 适合研究跨模态对齐、视觉-语言模型的开发者参考。
随着多模态学习的快速发展,图像-文本匹配作为连接视觉与语言的桥梁,日益重要。本文提出一种创新的视觉语义嵌入模型——多头共识感知视觉语义嵌入(MH-CVSE)。该模型在共识感知视觉语义嵌入(CVSE)基础上引入多头自注意力机制,可并行捕捉多个子空间中的信息,显著增强模型对图像与文本复杂关系的理解与表征能力。此外,采用参数化特征融合策略,灵活整合不同层次的特征信息,进一步提升模型表达力。在损失函数设计上,采用动态权重调整策略,根据损失值自动调节各损失项权重,使模型训练中更均衡地贡献不同损失项;同时引入余弦退火学习率策略,提升模型后期收敛稳定性。在Flickr30k数据集上的大量实验验证表明,MH-CVSE在双向图文检索任务中均优于先前方法,充分证明其有效性与优越性。
原文摘要 · Abstract (English)
With the rapid development of multimodal learning, the image-text matching task, as a bridge connecting vision and language, has become increasingly important. Based on existing research, this study proposes an innovative visual semantic embedding model, Multi-Headed Consensus-Aware Visual-Semantic Embedding (MH-CVSE). This model introduces a multi-head self-attention mechanism based on the consensus-aware visual semantic embedding model (CVSE) to capture information in multiple subspaces in parallel, significantly enhancing the model's ability to understand and represent the complex relationship between images and texts. In addition, we adopt a parameterized feature fusion strategy to flexibly integrate feature information at different levels, further improving the model's expressive power. In terms of loss function design, the MH-CVSE model adopts a dynamic weight adjustment strategy to dynamically adjust the weight according to the loss value itself, so that the model can better balance the contribution of different loss terms during training. At the same time, we introduce a cosine annealing learning rate strategy to help the model converge more stably in the later stages of training. Extensive experimental verification on the Flickr30k dataset shows that the MH-CVSE model achieves better performance than previous methods in both bidirectional image and text retrieval tasks, fully demonstrating its effectiveness and superiority.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。