arXiv:2607.14491cs.SIcs.CL2026-07

区分社交媒体操控中仇恨、党派攻击与地缘政治抨击,揭示误判根源。

Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations

论文配图:Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations
图 1 · 摘自论文原文
  • 用双提示大模型检测广义敌意,再通过可审计规则分类内容
  • 仅18.7%内容兼具身份攻击与去人性化,却占整体仇恨报告的近半
  • 六起操作归为三类,统一称为‘制造分裂’,适合研究网络操纵者

国家支持的社交平台影响力行动常被视作高发的‘仇恨’与‘毒性’来源。我们指出此类评估存在测量误差:现有检测器针对更广泛敌意(含对外群体攻击)验证,导致过度将党派或地缘政治攻击归为仇恨。分析推特信息操纵档案中七起政府关联活动(2508万条推文,8275个账号),我们首先用两提示大模型检测广义敌意(人类标注一致率κ=0.82),再构建可审计规则,对5457条内容进行三类细分:50.1%为针对个体的身份攻击,30.4%为党派攻击,19.5%为对国家及其外交政策的抨击。若全算作仇恨,会使其数量虚高近一倍;真正兼具身份攻击与去人性化或煽动性的仅占18.7%。六起行动可归为三类:身份仇恨(俄方,包括俄罗斯操作与IRA)、地缘政治抨击(两起伊朗操作)、党派分裂(两起委内瑞拉操作)。我们称其共同产物为“制造分裂”。在最难案例上,三位独立专家一致性仅为κ=0.37–0.50,十九个大模型最佳表现也仅达κ=0.601,说明界定边界仍存争议。本研究有助于重新定义影响行动中的仇恨概念及在线话语研究。

原文摘要 · Abstract (English)

State-backed influence operations are routinely measured as high-prevalence sources of ``hate'' and ``toxicity.'' We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definition inclusive of hostility or divisiveness aimed at an out-group, and so over-attribute hate to content better described as partisan or geopolitical invective. Across 25.08M tweets from seven government-attributed campaigns in the Twitter Information Operations archive (8,275 accounts), we separate hate from the other forms of divisiveness. We first validate a two-prompt LLM-based detector, matching human labels at Cohen's $κ=0.82$, to identify the broader hostility; we then develop an auditable rule, agreeing with an expert at $κ=0.52$, to further classify this content (5,457 posts) into three sub-categories. About 50.1% are identity-based attacks on people, whereas 30.4% are partisan attacks and 19.5% invective against states and their foreign policy. Reporting all of it as hate therefore overstates hate roughly twofold; only 18.7% is both identity-based and dehumanizing or inciting. Six of seven campaigns sort into three regimes that a single ``hate'' rate flattens, namely identity hate (RU-op and IRA, both Russia-attributed), geopolitical invective (both Iran operations), and partisan divisiveness (both Venezuela operations). We call the shared product $manufactured divisiveness$. The line to separate these constructs itself remains unsettled: on the hardest cases three independent human experts agree only moderately (pairwise $κ=0.37$--$0.50$), and the best of nineteen LLM models tops out at $κ=0.601$ against the experts' majority. Our findings can help redefine the study of hate in the context of influence campaigns and broader online discourse.

网络操纵仇恨检测语义分析地缘政治

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。