arXiv:2512.20061cs.AI2025-12被引 1

用强化学习提升大模型内容审核能力,数据效率比监督微调高100倍

Scaling Reinforcement Learning for Content Moderation with Large Language Models

  • 采用强化学习与可验证奖励机制,让大模型学会政策对齐的复杂推理
  • 在真实任务中表现优于监督微调,数据效率提升最高达100倍
  • 适合标注稀缺、政策多变的内容审核场景,对工业级系统有实用价值

大规模内容审核仍是当前数字生态中的核心挑战,需持续评估数十亿用户与AI生成内容是否违反政策。尽管大语言模型(LLMs)在政策对齐审核方面展现潜力,但如何在标签稀疏、政策动态变化、需深层推理的现实场景中训练出专家级准确率仍缺乏系统研究。本文全面实证研究了强化学习(RL)在内容分类中的扩展性,系统评估多种训练方案与奖励设计策略,包括可验证奖励和基于LLM的评判框架,将通用语言模型转化为三个真实任务中的专业化、政策对齐分类器。结果表明,RL呈现类S型扩展规律:性能随训练数据、采样轮次与优化步数平滑提升,最终趋于饱和;尤其在需要复杂政策推理的任务中,显著优于监督微调,数据效率最高提升100倍,特别适用于专家标注稀缺或成本高的领域。

原文摘要 · Abstract (English)

Content moderation at scale remains one of the most pressing challenges in today's digital ecosystem, where billions of user- and AI-generated artifacts must be continuously evaluated for policy violations. Although recent advances in large language models (LLMs) have demonstrated strong potential for policy-grounded moderation, the practical challenges of training these systems to achieve expert-level accuracy in real-world settings remain largely unexplored, particularly in regimes characterized by label sparsity, evolving policy definitions, and the need for nuanced reasoning beyond shallow pattern matching. In this work, we present a comprehensive empirical investigation of scaling reinforcement learning (RL) for content classification, systematically evaluating multiple RL training recipes and reward-shaping strategies-including verifiable rewards and LLM-as-judge frameworks-to transform general-purpose language models into specialized, policy-aligned classifiers across three real-world content moderation tasks. Our findings provide actionable insights for industrial-scale moderation systems, demonstrating that RL exhibits sigmoid-like scaling behavior in which performance improves smoothly with increased training data, rollouts, and optimization steps before gradually saturating. Moreover, we show that RL substantially improves performance on tasks requiring complex policy-grounded reasoning while achieving up to 100x higher data efficiency than supervised fine-tuning, making it particularly effective in domains where expert annotations are scarce or costly.

强化学习内容审核大模型应用数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。