用大模型模拟社交网络,评估不同内容审核策略的效果
Evaluating Online Moderation Via LLM-Powered Counterfactual Simulations
- 用大模型构建可交互的虚拟社交网络,模拟用户行为
- 发现个性化审核比统一规则更有效,且存在群体毒化传播现象
- 适合研究平台治理、算法伦理和内容安全的学者与工程师
在线社交网络(OSNs)广泛采用内容审核以遏制有害和攻击性言论的传播。然而,由于数据收集成本高且实验控制有限,现有审核干预的实际效果仍不明确。自然语言处理的最新进展为新评估方法提供了可能:大型语言模型(LLMs)可显著增强基于代理的建模,以极高的真实感模拟类人社交行为。然而,现有工具尚无法支持基于模拟的审核策略评估。本文填补该空白,设计了一个基于大模型的社交网络对话模拟器,实现并行反事实仿真——在保持其他条件不变的情况下,测试审核干预对毒性行为的影响。通过大量实验,揭示了社交代理的心理真实性、社会传染现象的出现,以及个性化审核策略的显著优势。
原文摘要 · Abstract (English)
Online Social Networks (OSNs) widely adopt content moderation to mitigate the spread of abusive and toxic discourse. Nonetheless, the real effectiveness of moderation interventions remains unclear due to the high cost of data collection and limited experimental control. The latest developments in Natural Language Processing pave the way for a new evaluation approach. Large Language Models (LLMs) can be successfully leveraged to enhance Agent-Based Modeling and simulate human-like social behavior with unprecedented degree of believability. Yet, existing tools do not support simulation-based evaluation of moderation strategies. We fill this gap by designing a LLM-powered simulator of OSN conversations enabling a parallel, counterfactual simulation where toxic behavior is influenced by moderation interventions, keeping all else equal. We conduct extensive experiments, unveiling the psychological realism of OSN agents, the emergence of social contagion phenomena and the superior effectiveness of personalized moderation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。