arXiv:2409.03219cs.CYcs.AI2024-09被引 56

LLM内容审核不应只看准确率,更要追求决策的正当性。

Content Moderation by LLM: From Accuracy to Legitimacy

  • 从准确率转向正当性框架,区分易判与难判案例
  • 难案需理由说明和用户参与,易案重速度透明
  • 适合关注平台治理、AI伦理的研究者与从业者

大语言模型(LLM)在在线平台内容审核中的应用日益广泛。现有研究多聚焦于准确率——即模型判断是否正确,但本文指出准确率不足以反映真实情况,因其忽视了易案与难案的区别,以及提升准确率必然伴随的权衡。内容审核实为平台治理的核心,关键在于获得并增强正当性。因此,本文主张从单一准确率评价转向以正当性为核心的评估框架:对易案,强调准确、快速与透明;对难案,则重视合理解释与用户参与。在此框架下,LLM真正潜力不在于提高准确率,而体现在四方面:区分难易案例、提供高质量解释、辅助人工审查获取上下文信息、促进更互动的用户参与。为此,文章提出一套整合流程,并运用法律与社会科学的规范理论批判性评估该技术应用,旨在重新定义LLM在内容审核中的角色,引导该领域研究方向转型。

原文摘要 · Abstract (English)

One trending application of LLM (large language model) is to use it for content moderation in online platforms. Most current studies on this application have focused on the metric of accuracy -- the extent to which LLMs make correct decisions about content. This article argues that accuracy is insufficient and misleading because it fails to grasp the distinction between easy cases and hard cases, as well as the inevitable trade-offs in achieving higher accuracy. Closer examination reveals that content moderation is a constitutive part of platform governance, the key of which is to gain and enhance legitimacy. Instead of making moderation decisions correct, the chief goal of LLMs is to make them legitimate. In this regard, this article proposes a paradigm shift from the single benchmark of accuracy towards a legitimacy-based framework for evaluating the performance of LLM moderators. The framework suggests that for easy cases, the key is to ensure accuracy, speed, and transparency, while for hard cases, what matters is reasoned justification and user participation. Examined under this framework, LLMs' real potential in moderation is not accuracy improvement. Rather, LLMs can better contribute in four other aspects: to conduct screening of hard cases from easy cases, to provide quality explanations for moderation decisions, to assist human reviewers in getting more contextual information, and to facilitate user participation in a more interactive way. To realize these contributions, this article proposes a workflow for incorporating LLMs into the content moderation system. Using normative theories from law and social sciences to critically assess the new technological application, this article seeks to redefine LLMs' role in content moderation and redirect relevant research in this field.

内容审核正当性LLM治理平台治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。