arXiv:2609.03814cs.CL2026-09

测试大模型在内容审核中是否能区分不同判断标准,发现高分背后隐藏问题。

Evaluating Criterion-Conditioned Behaviour of Large Language Models in Content Moderation

  • 设计新评估方法DECO,分离内容特征实现按标准独立测试
  • 4个模型在4个数据集上表现差异大,特定标准下准确率低至50%
  • 适合关注审核公平性与模型可解释性的研究者和工程师

大型语言模型(LLMs)在标准内容审核基准上表现优异,但这些基准常将多个审核标准合并为单一标签,难以判断模型能否区分并准确应用各标准。为研究模型是否具备准则条件行为,我们提出诊断性内容评估(DECO),一种独立于准则的内容因子化方法,支持在控制条件下进行准则级评估,并引入成对评估比较同一输入在不同准则下的模型输出。在四个审核数据集和四个LLM上的实验表明,尽管基准整体表现良好,但在具体准则层面存在显著失败。当正确决策不依赖于整体有害性,而取决于特定准则要求的方面时,模型表现最差。结果凸显了当前审核基准的关键局限:在聚合标签上表现良好,不足以证明模型能可靠地针对个别准则评估内容。研究呼吁开发明确测量准则条件行为的评估方法。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate strong performance on standard content moderation benchmarks. However, these benchmarks often aggregate multiple moderation criteria into a single label, making it unclear whether models can disentangle them and reliably apply each criterion when making decisions. To study whether LLMs exhibit criterion-conditioned behaviour, we introduce Diagnostic Evaluation of COntent (DECO), a criterion-independent factorisation of content that enables controlled, criterion-level evaluation. We also introduce pairwise evaluation to compare model outputs across different criteria for the same input. Across four moderation datasets and four LLMs, we find that strong benchmark performance can hide substantial failures at the criterion level. Models struggle most when correct decisions depend not on overall harmfulness, but on the specific aspect of the content that the criterion requires them to assess. Our results highlight a key limitation of current content moderation benchmarks: strong performance on aggregated labels does not provide sufficient evidence that LLMs can reliably evaluate content with respect to individual moderation criteria. These findings call for the development of evaluation methods that explicitly measure criterion-conditioned behaviour.

内容审核评测方法大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。