arXiv:2409.08963cs.CYcs.CL2024-09被引 8

用大模型代理自动检测去中心化社交平台违规内容,效果优于人工审核。

Safeguarding Decentralized Social Media: LLM Agents for Automating Community Rule Compliance

  • 基于开源大模型构建六种AI代理,自动判断社区规则合规性。
  • 在超5万条Mastodon帖子中,识别准确率高且能理解语言细节。
  • 适合需要半自动或人机协同的内容审核系统使用。

确保内容符合社区准则对维护健康的在线社交环境至关重要。然而,传统依赖人工的合规检查因用户生成内容激增和审核员有限而难以扩展。近年来,大型语言模型在自然语言理解上的进展为自动化内容合规验证提供了新可能。本研究评估了六个基于Open-LLMs构建的AI代理,在去中心化社交网络中自动执行规则合规检查的能力。该环境挑战在于社区范围与规则高度异质。通过对数百个Mastodon服务器的超过5万条帖子进行分析,发现这些AI代理能有效检测违规内容,理解语言细微差别,并适应多样化的社区语境。多数代理在评分解释和合规建议上表现出高一致性与跨评估者可靠性。领域专家的人工评估确认了其可靠性与实用性,表明其是半自动或人机协同内容审核系统的有力工具。

原文摘要 · Abstract (English)

Ensuring content compliance with community guidelines is crucial for maintaining healthy online social environments. However, traditional human-based compliance checking struggles with scaling due to the increasing volume of user-generated content and a limited number of moderators. Recent advancements in Natural Language Understanding demonstrated by Large Language Models unlock new opportunities for automated content compliance verification. This work evaluates six AI-agents built on Open-LLMs for automated rule compliance checking in Decentralized Social Networks, a challenging environment due to heterogeneous community scopes and rules. Analyzing over 50,000 posts from hundreds of Mastodon servers, we find that AI-agents effectively detect non-compliant content, grasp linguistic subtleties, and adapt to diverse community contexts. Most agents also show high inter-rater reliability and consistency in score justification and suggestions for compliance. Human-based evaluation with domain experts confirmed the agents' reliability and usefulness, rendering them promising tools for semi-automated or human-in-the-loop content moderation systems.

AI代理内容审核去中心化社交大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。