arXiv:2409.14552cs.CLcs.AI2024-09EMNLP

用图模型联合建模文本与表情符号,提升社交媒体分析效果

Unleashing the Power of Emojis in Texts via Self-supervised Graph Pre-Training

  • 构建包含帖子、词语和表情符号的异构图,显式建模三者交互
  • 通过节点对比学习和边重建任务,在小红书和推特数据集上显著提效
  • 适合做社交内容理解、情感分析的研究者与工程师参考

表情符号在社交平台上广受欢迎,常用于补充或替代文字。然而,现有数据挖掘方法通常忽略或简单将表情符号当作普通字符处理,限制了模型对表情符号语义及与文本互动的理解。为此,本文构建一个包含帖子、词语和表情符号三类节点的异构图,并明确定义它们之间的边关系,以增强内容表征。为促进三类节点间信息共享,提出一种图文协同预训练框架,包含两个图预训练任务:节点级对比学习与边级链接重建学习。在小红书和推特数据集上的两类下游任务实验表明,该方法显著优于先前强基线模型。

原文摘要 · Abstract (English)

Emojis have gained immense popularity on social platforms, serving as a common means to supplement or replace text. However, existing data mining approaches generally either completely ignore or simply treat emojis as ordinary Unicode characters, which may limit the model's ability to grasp the rich semantic information in emojis and the interaction between emojis and texts. Thus, it is necessary to release the emoji's power in social media data mining. To this end, we first construct a heterogeneous graph consisting of three types of nodes, i.e. post, word and emoji nodes to improve the representation of different elements in posts. The edges are also well-defined to model how these three elements interact with each other. To facilitate the sharing of information among post, word and emoji nodes, we propose a graph pre-train framework for text and emoji co-modeling, which contains two graph pre-training tasks: node-level graph contrastive learning and edge-level link reconstruction learning. Extensive experiments on the Xiaohongshu and Twitter datasets with two types of downstream tasks demonstrate that our approach proves significant improvement over previous strong baseline methods.

图神经网络表情符号多模态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。