arXiv:2412.03930cs.CLcs.AI2024-12KDD被引 7

用图文结合方法提升文本图谱中的异常检测精度与效率

GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language Model

  • 融合图结构特征与小模型提取的语义属性
  • 在4个数据集上超越现有方法,训练/推理提速5倍
  • 适合需要高效检测社交网络、学术关联异常的场景

文本丰富的图谱中异常检测广泛存在于现实场景,如错误归属论文给作者、社交网络中的机器人识别。大语言模型(LLM)利用丰富文本信息为异常检测开辟新路径,但直接引入文本会掩盖关键检测线索并带来高昂微调成本。此外,LLM常忽略图的内在结构偏置,而该偏置对区分正常与异常节点模式至关重要。为此,本文提出GuARD,一种结合图方法关键结构特征与小语言模型提取的细粒度语义属性的文本丰富且图感知语言模型,用于文本丰富图谱的异常检测。GuARD采用任务引导的渐进式多模态多轮指令微调框架进行优化,以融合丰富文本与结构模态。在四个数据集上的大量实验表明,GuARD优于基于图和基于LLM的异常检测方法,且在大规模WhoIsWho数据集上相较原始长上下文LLM实现高达5倍的训练加速和5倍的推理加速。

原文摘要 · Abstract (English)

Anomaly detection on text-rich graphs is widely prevalent in real life, such as detecting incorrectly assigned academic papers to authors and detecting bots in social networks. The remarkable capabilities of large language models (LLMs) pave a new revenue by utilizing rich-text information for effective anomaly detection. However, simply introducing rich texts into LLMs can obscure essential detection cues and introduce high fine-tuning costs. Moreover, LLMs often overlook the intrinsic structural bias of graphs which is vital for distinguishing normal from abnormal node patterns. To this end, this paper introduces GuARD, a text-rich and graph-informed language model that combines key structural features from graph-based methods with fine-grained semantic attributes extracted via small language models for effective anomaly detection on text-rich graphs. GuARD is optimized with the progressive multi-modal multi-turn instruction tuning framework in the task-guided instruction tuning regime tailed to incorporate both rich-text and structural modalities. Extensive experiments on four datasets reveal that GuARD outperforms graph-based and LLM-based anomaly detection methods, while offering up to 5$\times$ times speedup in training and 5$\times$ times speedup in inference over vanilla long-context LLMs on the large-scale WhoIsWho dataset.

异常检测图神经网络大模型应用文本图谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。