用AI辅助制定详细分类规则,提升内容审核标签一致性。
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation

- AI生成细粒度分类宪法,覆盖边界案例
- 相比段落定义,模型间不一致降低57倍
- 适合需要高一致性的内容审核场景
许多自动化标注流程将输入分类到由书面规范定义的类别中,内容审核是典型应用。简单的类别定义不足以让标注者生成准确、一致的黄金标签。一种解决方案是编写详尽的规范性定义,明确解决足够多的实际边界案例,使标注者无法对解释产生分歧。但实践中,这种细节程度超出人类工作记忆容量,标注者只能依赖直觉,导致标签偏离规则,准确性和一致性下降。我们提出并验证了一种AI驱动的工作流:AI协助制定每个类别的详细宪法,以涵盖边缘情况;前沿大模型据此解析每条输入,生成比人类更一致、更准确的黄金标签。我们在三个内容审核类别(骚扰、仇恨言论、非暴力犯罪)上评估,结果显示该方法将跨模型不一致降低至段落定义的1/57,且模型间分歧可诊断出规范漏洞。人类仅负责高层次定义决策,而非具体标注判断。为安全评估引入双轴评分机制,独立评估意图与内容,下游使用者可选择作用于任一轴或两者。
原文摘要 · Abstract (English)
Many automated labeling pipelines classify inputs into categories defined by a written specification, content moderation being a prominent use case. Simple category definitions are not detailed enough for labelers to produce the accurate, consistent golden labels these pipelines require. One solution is to write a prescriptive definition that settles enough real boundary cases that labelers cannot disagree with the written interpretation. In practice, definitions at that level of detail exceed what a human annotator can hold in working memory, so annotators fall back on intuition and the labels drift from the written rules, regressing on accuracy and consistency. We propose and demonstrate the efficacy of an AI-driven workflow in which AI helps write a per-category constitution that defines the label in enough detail to cover edge cases, and a frontier LLM interprets it on each input to produce the golden label more consistently and accurately than humans reading the same document. We evaluate on three content moderation categories (harassment, hate speech, non-violent crime) and show that the approach reduces cross-model inconsistency by up to 57x compared to paragraph definitions, with cross-model disagreement diagnosing specification gaps and the human responsible for high-level decisions about what each category should mean rather than individual labeling calls. For the safety evaluation, we introduce a dual-axis formulation scoring intent and content independently over the full conversation, so downstream consumers can act on either axis or both.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。