用可解释的AI识别国家支持的网络水军,还能说明他们用了什么话术。
X-Troll: eXplainable Detection of State-Sponsored Information Operations Agents
- 用领域知识增强LLM,通过特殊适配器捕捉操纵性语言特征。
- 在真实数据上准确率超越普通大模型和现有检测方法。
- 输出人类能看懂的解释,揭示水军具体使用的宣传策略。
国家支持的网络水军通过有组织的信息操控行为,对在线话语生态构成威胁。尽管大语言模型(LLM)在通用自然语言处理任务中表现优异,但在识别微妙的宣传手段方面能力有限,且运行过程如同“黑箱”,无法提供可解释的决策依据。本文提出X-Troll框架,将可解释的适配器增强型LLM与专家级语言学知识结合,实现对国家支持水军的检测,并生成人类可读的决策解释。X-Troll融合评估理论与宣传分析,通过专用LoRA适配器动态捕捉协同信息操作中的特定话语模式。在真实数据上的实验表明,该方法在准确率上优于通用大模型和现有水军检测模型,同时通过基于专家知识的解释,揭示了国家支持行为所采用的具体语言策略。代码已开源:https://github.com/ltian678/xtroll_source/
原文摘要 · Abstract (English)
State-sponsored trolls, malicious actors who deploy sophisticated linguistic manipulation in coordinated information campaigns, posing threats to online discourse integrity. While Large Language Models (LLMs) achieve strong performance on general natural language processing (NLP) tasks, they struggle with subtle propaganda detection and operate as ``black boxes'', providing no interpretable insights into manipulation strategies. This paper introduces X-Troll, a novel framework that bridges this gap by integrating explainable adapter-based LLMs with expert-derived linguistic knowledge to detect state-sponsored trolls and provide human-readable explanations for its decisions. X-Troll incorporates appraisal theory and propaganda analysis through specialized LoRA adapters, using dynamic gating to capture campaign-specific discourse patterns in coordinated information operations. Experiments on real-world data demonstrate that our linguistically-informed approach shows strong performance compared with both general LLM baselines and existing troll detection models in accuracy while providing enhanced transparency through expert-grounded explanations that reveal the specific linguistic strategies used by state-sponsored actors. X-Troll source code is available at: https://github.com/ltian678/xtroll_source/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。