用标注者判断诊断主观任务标注方案,找出分歧根源。
Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks

- 基于多标注者判断,分析标注准则的稳定性与重叠性。
- 发现少数准则不稳,近半句子跨多个类别。
- 适合改进标注指南或调整分类结构的团队使用。
主观NLP数据集通常将标注者意见合并为单一黄金标签,难以判断分歧是因标准模糊、区分失效,还是本就存在合理多样性。本文提出一种在确定黄金标签前对专家设计的标注方案进行评估的「模式级诊断」方法,仅需多标注者的准则判断。该方法可区分两种失败模式:准则边界难以操作化(不稳定)和类别间系统性重叠(边界模糊)。应用于商业文件中的说服价值提取任务时,发现分歧并非随机分布:少数准则存在不稳定性,近半数句子同时激活多个类别。这些信号与领域专家的分歧一致,为优化标注指南、重构类别体系或重新考虑标注范式提供了实证依据。
原文摘要 · Abstract (English)
Subjective NLP datasets typically aggregate annotator judgments into a single gold label, making it difficult to diagnose whether disagreement reflects unclear criteria, collapsed distinctions, or legitimate plurality. We propose a \emph{schema-level diagnostic} for auditing expert-designed annotation schemas \emph{prior to} gold-label commitment, using only multi-annotator criterion judgments. The diagnostic separates two failure modes: unstable criteria with hard-to-operationalize boundaries, and systematic overlap that blurs the boundaries between mutually exclusive categories. Applied to persuasive value extraction in commercial documents, we find that disagreement is not diffuse: instability concentrates in a few criteria, while nearly half of covered sentences activate multiple categories. These signals align with where domain experts disagree, yielding an evidence-based audit for tightening guidelines, revising category structure, or reconsidering the annotation paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。