用多LLM代理模拟仇恨言论传播,更真实还原其扩散模式。
Simulating Hate Speech Cascades with Multi-LLM Agents: Empirical Grounding, Modeling Fidelity, and Intervention Strategies
- 构建多LLM代理系统,让转发决策受用户、社区和内容影响。
- 模拟结果再现97.4%以上转发者持敌对立场,且扩散呈星型结构。
- 发现代理差异性是模拟关键,针对性干预可降噪7.5%-12.9%。
在线平台上的仇恨内容传播建模仍是监管研究中的开放问题。传统级联模型未显式刻画用户特征、社区环境与内容因素,导致实际部署时策略效果不佳。多智能体大语言模型系统理论上可使每次转发决策依赖用户画像、社群背景与内容本身,但其是否比经典基线更真实还原真实仇恨级联尚不明确。我们研究了三个仇恨类蓝溪(Bluesky)级联及一个规模匹配的良性对照。在实证数据中发现:97.4–99.7%的转发者持敌意立场;仇恨级联中毒性-参与同质性在扩散树上高于关注图;仇恨级联呈星型拓扑(多数转发直接来自源头),而良性级联为树状(经多跳传播)。模拟结果显示,多LLM代理模拟器成功复现了立场单一化与毒性-增量方向。结构化消融分析表明,代理异质性是提升拟合度的核心因素,针对密集网络的放大器干预可在5.7%良性误伤下实现7.5–12.9%的传播抑制。
原文摘要 · Abstract (English)
Faithful modeling of hateful content propagation on online platforms remains an open problem for moderation research. Classical cascade models that do not explicitly represent the profile, community, and content factors associated with hateful-content propagation may yield moderation strategies that behave less effectively when deployed in real-world scenarios. Multi-agent large language model (LLM) systems can, in principle, make each reshare decision depend on the user's profile, the surrounding community, and the post's content, but it remains unclear whether this added flexibility actually reproduces real hateful cascades more faithfully than classical baselines. We study three hateful Bluesky cascades and a size-matched benign control. In the empirical Bluesky data, we found that: 97.4--99.7\% of reposters take a hostile stance; toxicity-engagement homophily is higher on the diffusion tree than on the follower graph for hateful cascades; topology is star-like for the hateful cascades (most reposts come directly from the root) versus tree-like for the benign cascade (reposts propagate through multi-hop chains). In simulation, a multi-LLM-agent simulator reproduces the stance monoculture and the toxicity-delta direction. A structured ablation identifies agent heterogeneity as the leading fidelity factor, and amplifier targeting on dense networks yields 7.5--12.9\% reduction at 5.7\% benign collateral.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。