arXiv:2507.20061cs.LGcs.GT2025-07被引 1

设计智能内容过滤机制,在保护言论自由与减少操纵行为间找平衡。

Strategic Filtering for Content Moderation: Free Speech or Free of Distortion?

  • 从机制设计视角优化过滤策略,兼顾言论自由与社会失真
  • 提出可近似最优解的实用方法,证明有限离线数据即可有效逼近
  • 适合研究平台治理、算法伦理与内容安全的学者和工程师

社交媒体上的用户生成内容易受煽动与操纵影响,亟需有效监管。平台常部署自动化内容审核系统,评估内容危害性并过滤违规信息。但此类审核会引发用户的策略性应对,即在规则限制下表达观点。这要求在两个目标间取得平衡:1)保障言论自由——最小化表达限制;2)降低社会失真——以总内容操纵量衡量。本文从机制设计角度解决该权衡问题,旨在最小化社会失真同时最大化言论自由。尽管最优权衡判定为NP难,我们提出了可行的近似方法,并提供了泛化保证,确定了实现有效近似所需的有限离线数据量。

原文摘要 · Abstract (English)

User-generated content (UGC) on social media platforms is vulnerable to incitements and manipulations, necessitating effective regulations. To address these challenges, those platforms often deploy automated content moderators tasked with evaluating the harmfulness of UGC and filtering out content that violates established guidelines. However, such moderation inevitably gives rise to strategic responses from users, who strive to express themselves within the confines of guidelines. Such phenomena call for a careful balance between: 1. ensuring freedom of speech -- by minimizing the restriction of expression; and 2. reducing social distortion -- measured by the total amount of content manipulation. We tackle the problem of optimizing this balance through the lens of mechanism design, aiming at optimizing the trade-off between minimizing social distortion and maximizing free speech. Although determining the optimal trade-off is NP-hard, we propose practical methods to approximate the optimal solution. Additionally, we provide generalization guarantees determining the amount of finite offline data required to approximate the optimal moderator effectively.

内容审核机制设计言论自由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。