arXiv:2502.18695cs.CYcs.AI2025-02被引 44

用自然语言直接让大模型执行内容审核,省去数据标注环节。

Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models

  • 将政策文本转为提示词,让大模型直接理解并执行审核
  • 发现提示结构变化会显著影响审核结果的稳定性
  • 适合平台安全团队与AI治理研究者参考

内容审核对构建安全包容的网络环境至关重要,需平衡平台规范、用户期待与监管要求。传统方式依赖将政策转化为指南,供人工审核或用于训练机器学习模型。但大语言模型的发展使政策可直接作为文本输入进行解析,无需大量数据标注。这一新范式带来动态调整审核策略的可能,也引发新问题:政策如何有效转化?我们正式提出“政策即提示”框架,识别出四大领域中的五大挑战——技术实现(1. 政策转提示,2. 对提示结构敏感)、社会技术(3. 技术决定政策形成风险)、组织管理(4. 政策与算法团队角色演变)、治理机制(5. 模型治理与问责)。通过多维度分析,提出缓解路径,为可扩展、自适应的内容审核系统提供实践指引。

原文摘要 · Abstract (English)

Content moderation plays a critical role in shaping safe and inclusive online environments, balancing platform standards, user expectations, and regulatory frameworks. Traditionally, this process involves operationalising policies into guidelines, which are then used by downstream human moderators for enforcement, or to further annotate datasets for training machine learning moderation models. However, recent advancements in large language models (LLMs) are transforming this landscape. These models can now interpret policies directly as textual inputs, eliminating the need for extensive data curation. This approach offers unprecedented flexibility, as moderation can be dynamically adjusted through natural language interactions. This paradigm shift raises important questions about how policies are operationalised and the implications for content moderation practices. In this paper, we formalise the emerging policy-as-prompt framework and identify five key challenges across four domains: Technical Implementation (1. translating policy to prompts, 2. sensitivity to prompt structure and formatting), Sociotechnical (3. the risk of technological determinism in policy formation), Organisational (4. evolving roles between policy and machine learning teams), and Governance (5. model governance and accountability). Through analysing these challenges across technical, sociotechnical, organisational, and governance dimensions, we discuss potential mitigation approaches. This research provides actionable insights for practitioners and lays the groundwork for future exploration of scalable and adaptive content moderation systems in digital ecosystems.

内容审核大模型应用AI治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。