利用多源网络知识提升无标签场景下的社交机器人检测效果
BotTrans: A Multi-Source Graph Domain Adaptation Approach for Social Bot Detection
- 构建跨源图结构增强同质性,融合多源邻居信息提升特征判别力
- 通过源-目标相关性加权优化,优先迁移与任务更相关的知识
- 适合需要在无标注数据上检测社交机器人的人工智能安全研究者
基于图神经网络的社交机器人检测面临标签稀缺问题。现有方法依赖单一源域迁移,易受网络异质性影响且知识有限。本文提出多源图域适应模型BotTrans:首先利用多源网络共享标签构建跨域同质拓扑,增强源节点表征;再聚合跨域邻域信息以提升特征区分能力;随后在模型优化中引入源-目标相关性权重,优先转移相关性强的源域知识;最后设计细粒度优化策略,利用目标域语义信息进一步提升性能。在真实数据集上的实验表明,BotTrans显著优于现有最先进方法,在无标签目标域上实现更稳定、更准确的检测。
原文摘要 · Abstract (English)
Transferring extensive knowledge from relevant social networks has emerged as a promising solution to overcome label scarcity in detecting social bots and other anomalies with GNN-based models. However, effective transfer faces two critical challenges. Firstly, the network heterophily problem, which is caused by bots hiding malicious behaviors via indiscriminately interacting with human users, hinders the model's ability to learn sufficient and accurate bot-related knowledge from source domains. Secondly, single-source transfer might lead to inferior and unstable results, as the source network may embody weak relevance to the task and provide limited knowledge. To address these challenges, we explore multiple source domains and propose a multi-source graph domain adaptation model named \textit{BotTrans}. We initially leverage the labeling knowledge shared across multiple source networks to establish a cross-source-domain topology with increased network homophily. We then aggregate cross-domain neighbor information to enhance the discriminability of source node embeddings. Subsequently, we integrate the relevance between each source-target pair with model optimization, which facilitates knowledge transfer from source networks that are more relevant to the detection task. Additionally, we propose a refinement strategy to improve detection performance by utilizing semantic knowledge within the target domain. Extensive experiments on real-world datasets demonstrate that \textit{BotTrans} outperforms the existing state-of-the-art methods, revealing its efficacy in leveraging multi-source knowledge when the target detection task is unlabeled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。