arXiv:2501.08267cs.IRcs.SI2025-01中稿 · CASCON被引 5

融合文本、图像和标签三模态信息,提升社交媒体命名实体识别准确率。

TriMod Fusion for Multimodal Named Entity Recognition in Social Media

  • 设计三模态融合架构,用Transformer注意力整合文本、图像与标签特征。
  • 在多模态社交媒体数据集上,F1分数显著优于现有方法。
  • 适合需要精准理解社交语境的多模态分析任务使用。

社交媒体平台是用户生成内容的重要来源,为理解人类行为提供了丰富信息。命名实体识别(NER)通过识别并分类命名实体,对这类内容分析至关重要。然而,传统NER模型难以应对社交媒体语言中常见的非正式、上下文稀疏和语义模糊问题。近期研究转向多模态方法,结合文本与视觉线索以提升识别效果。尽管如此,现有方法在捕捉视觉对象与文本实体间的细微关联,以及解决模态间分布差异方面仍存在局限。本文提出一种新方法,融合文本、视觉和标签三类特征(TriMod),利用Transformer注意力实现高效模态融合。实验表明,该模型在多模态社交媒体数据集上显著优于现有先进方法,在精确率、召回率和F1分数上均有提升,证明多模态辅助上下文能极大增强命名实体识别能力。

原文摘要 · Abstract (English)

Social media platforms serve as invaluable sources of user-generated content, offering insights into various aspects of human behavior. Named Entity Recognition (NER) plays a crucial role in analyzing such content by identifying and categorizing named entities into predefined classes. However, traditional NER models often struggle with the informal, contextually sparse, and ambiguous nature of social media language. To address these challenges, recent research has focused on multimodal approaches that leverage both textual and visual cues for enhanced entity recognition. Despite advances, existing methods face limitations in capturing nuanced mappings between visual objects and textual entities and addressing distributional disparities between modalities. In this paper, we propose a novel approach that integrates textual, visual, and hashtag features (TriMod), utilizing Transformer-attention for effective modality fusion. The improvements exhibited by our model suggest that named entities can greatly benefit from the auxiliary context provided by multiple modalities, enabling more accurate recognition. Through the experiments on a multimodal social media dataset, we demonstrate the superiority of our approach over existing state-of-the-art methods, achieving significant improvements in precision, recall, and F1 score.

多模态命名实体识别社交媒体Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。