提出融合文本与图结构的新型异常检测框架,提升社交网络中真假用户识别能力。
Text Meets Topology: Rethinking Out-of-distribution Detection in Text-Rich Networks
- 设计跨注意力模块融合局部图结构到文本表示
- 通过超网络生成节点专属变换参数,增强特征区分度
- 在11个数据集上验证多场景下检测性能,适合复杂网络分析
文本丰富网络中的分布外(OOD)检测仍具挑战,因文本特征与拓扑结构交织。现有方法仅处理标签偏移或简单领域划分,忽略文本与结构的复杂多样性。例如,在社交网络中,用户作为节点具有文本属性(名称、简介),边表示好友关系,异常可能源于机器人与正常用户间的语言模式差异。为此,我们提出TextTopoOOD框架,评估四类典型OOD场景:(1)通过文本增强和嵌入扰动实现属性级偏移;(2)通过边重连和语义连接引入结构偏移;(3)主题引导的标签偏移;(4)基于领域的划分。同时提出TNT-OOD模型,利用新型交叉注意力模块将局部结构融入节点文本表示,并采用超网络生成节点特异的变换参数,使同类节点在拓扑与语义上对齐,从而在结构与文本偏移下增强ID与OOD的区分能力。在11个数据集上的实验验证了TextTopoOOD在多场景下的评估价值。
原文摘要 · Abstract (English)
Out-of-distribution (OOD) detection remains challenging in text-rich networks, where textual features intertwine with topological structures. Existing methods primarily address label shifts or rudimentary domain-based splits, overlooking the intricate textual-structural diversity. For example, in social networks, where users represent nodes with textual features (name, bio) while edges indicate friendship status, OOD may stem from the distinct language patterns between bot and normal users. To address this gap, we introduce the TextTopoOOD framework for evaluating detection across diverse OOD scenarios: (1) attribute-level shifts via text augmentations and embedding perturbations; (2) structural shifts through edge rewiring and semantic connections; (3) thematically-guided label shifts; and (4) domain-based divisions. Furthermore, we propose TNT-OOD to model the complex interplay between Text aNd Topology using: 1) a novel cross-attention module to fuse local structure into node-level text representations, and 2) a HyperNetwork to generate node-specific transformation parameters. This aligns topological and semantic features of ID nodes, enhancing ID/OOD distinction across structural and textual shifts. Experiments on 11 datasets across four OOD scenarios demonstrate the nuanced challenge of TextTopoOOD for evaluating OOD detection in text-rich networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。