arXiv:2412.09870cs.CV2024-12

通过动态对齐视觉与文本特征,提升社交媒体定位的准确性。

Dynamic Cross-Modal Alignment for Robust Semantic Location Prediction

  • 引入上下文感知的视觉语言对齐模块,增强跨模态特征匹配。
  • 在基准数据集上准确率提升2.3%,F1分数提升2.5%。
  • 对噪声鲁棒性强,适合真实场景下的位置预测任务。

从多模态社交媒体内容中进行语义位置预测是一项关键任务,广泛应用于个性化服务与人类移动性分析。本文提出一种判别式框架——上下文感知视觉语言对齐(CoVLA),以应对该任务中存在的语境歧义与模态差异问题。CoVLA利用上下文对齐模块(CAM)增强跨模态特征对齐,并通过跨模态融合模块(CMF)动态整合文本与视觉信息。在基准数据集上的大量实验表明,CoVLA显著优于现有最优方法,准确率提升2.3%,F1分数提升2.5%。消融实验证明了CAM与CMF的有效性,人工评估也验证了预测结果的上下文相关性。此外,鲁棒性分析显示,即使在噪声环境下,CoVLA仍保持高性能,展现出在真实应用中的可靠性。这些结果凸显了CoVLA在推动语义位置预测研究方面的潜力。

原文摘要 · Abstract (English)

Semantic location prediction from multimodal social media posts is a critical task with applications in personalized services and human mobility analysis. This paper introduces \textit{Contextualized Vision-Language Alignment (CoVLA)}, a discriminative framework designed to address the challenges of contextual ambiguity and modality discrepancy inherent in this task. CoVLA leverages a Contextual Alignment Module (CAM) to enhance cross-modal feature alignment and a Cross-modal Fusion Module (CMF) to dynamically integrate textual and visual information. Extensive experiments on a benchmark dataset demonstrate that CoVLA significantly outperforms state-of-the-art methods, achieving improvements of 2.3\% in accuracy and 2.5\% in F1-score. Ablation studies validate the contributions of CAM and CMF, while human evaluations highlight the contextual relevance of the predictions. Additionally, robustness analysis shows that CoVLA maintains high performance under noisy conditions, making it a reliable solution for real-world applications. These results underscore the potential of CoVLA in advancing semantic location prediction research.

多模态位置预测视觉语言对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。