arXiv:2603.01724cs.AI2026-03

构建新基准评估AI在多重违规与规则变化下的内容审核能力

GMP: A Benchmark for Content Moderation under Co-occurring Violations and Dynamic Rules

  • 提出GMP基准,模拟真实场景中多重违规与动态规则
  • 揭示现有模型在规则变化时判断力显著下降
  • 适合研究可信AI、内容审核与评测方法的学者

在线内容审核对维护健康数字环境至关重要,其对AI的依赖持续增长。现实场景中,一条评论可能同时违反多个政策(如歧视性言论与人身攻击),且平台规则随上下文动态变化。当前大语言模型虽能遵循固定规则,但在政策不稳或情境依赖时判断能力大幅下降,导致审核结果不一致:要么误删合法内容,要么放任有害信息。这引发关键问题:现有静态基准上的高表现是否真能保证模型在复杂真实场景中的泛化能力?为此,本文提出GMP基准,专门评估AI在多重违规共存与规则动态变化下的内容审核表现。

原文摘要 · Abstract (English)

Online content moderation is essential for maintaining a healthy digital environment, and reliance on AI for this task continues to grow. Consider a user comment using national stereotypes to insult a politician. This example illustrates two critical challenges in real-world scenarios: (1) Co-occurring Violations, where a single post violates multiple policies (e.g., prejudice and personal attacks); (2) Dynamic rules of moderation, where determination of a violation depends on platform-specific guidelines that evolve across contexts . The intersection of co-occurring harms and dynamically changing rules highlights a core limitation of current AI systems: although large language models (LLMs) are adept at following fixed guidelines, their judgment capabilities degrade when policies are unstable or context-dependent . In practice, such shortcomings lead to inconsistent moderation: either erroneously restricting legitimate expression or allowing harmful content to remain online . This raises a critical question for evaluation: Does high performance on existing static benchmarks truly guarantee robust generalization of AI judgment to real-world scenarios involving co-occurring violations and dynamically changing rules?

内容审核多任务动态规则评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。