arXiv:2605.01416cs.CYcs.CL2026-05中稿 · the 34th European …

用多智能体框架让内容审核更符合个人敏感度,提升判断精准度。

Who Decides What Is Harmful? Content Moderation Policy Through A Multi-Agent Personalised Inference Framework

  • 基于大模型构建多智能体系统,按用户敏感度个性化过滤内容。
  • 相比传统方法,准确率最高提升32%,更贴近个体感知。
  • 适合关注平台治理与数字权利平衡的研究者和从业者。

在线平台规模与复杂性的增长带来了有害内容、数字健康与用户自主权的关键政策问题。传统内容审核依赖集中式、自上而下的规则,难以适应伤害感知的主观性。本文提出一种基于大语言模型的多智能体个性化推理框架,根据用户独特的敏感度特征过滤内容。该架构包含领域专家智能体、负责内容分析与智能体调度的管理智能体,以及模拟用户视角的幽灵画像智能体,以支持审核决策。在多种非个性化基线对比下,系统表现出最高达32%的准确率提升,显著增强与个体敏感度的一致性。除技术性能外,该框架为平台治理提供了政策相关洞见,提供了一种可扩展的方式,以调和内容审核政策与社会及个体数字权利之间的矛盾。

原文摘要 · Abstract (English)

The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy. Traditional content moderation systems rely on centralised, top-down rules, often failing to accommodate the subjective nature of harm perception. This paper proposes an LLM-based multi-agent personalised inference framework that filters content based on unique sensitivity profiles of individual users. Our architecture combines domain-specific Expert Agents, a Manager Agent for orchestrating content analysis and agent selection, and a Ghost Profile Agent for simulating user perspectives, to inform moderation decisions. Evaluated against a range of non-personalised baselines, the system demonstrates up to a 32% improvement in accuracy, showing increased alignment with individual user sensitivities. Beyond technical performance, our framework provides policy-relevant insights for platform governance, providing a scalable way to reconcile moderation policies with societal and individual digital rights

内容审核多智能体个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。