arXiv:2604.17022cs.CLcs.AI2026-04ACL

用标注者判断诊断主观任务标注方案,找出分歧根源。

Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks

论文配图:Beyond Black-Box Labels: Interpretable Criteria for Diagnosing Subjective NLP Tasks
图 1 · 摘自论文原文
  • 基于多标注者判断,分析标注准则的稳定性与重叠性。
  • 发现少数准则不稳,近半句子跨多个类别。
  • 适合改进标注指南或调整分类结构的团队使用。

主观NLP数据集通常将标注者意见合并为单一黄金标签,难以判断分歧是因标准模糊、区分失效,还是本就存在合理多样性。本文提出一种在确定黄金标签前对专家设计的标注方案进行评估的「模式级诊断」方法,仅需多标注者的准则判断。该方法可区分两种失败模式:准则边界难以操作化(不稳定)和类别间系统性重叠(边界模糊)。应用于商业文件中的说服价值提取任务时,发现分歧并非随机分布:少数准则存在不稳定性,近半数句子同时激活多个类别。这些信号与领域专家的分歧一致,为优化标注指南、重构类别体系或重新考虑标注范式提供了实证依据。

原文摘要 · Abstract (English)

Subjective NLP datasets typically aggregate annotator judgments into a single gold label, making it difficult to diagnose whether disagreement reflects unclear criteria, collapsed distinctions, or legitimate plurality. We propose a \emph{schema-level diagnostic} for auditing expert-designed annotation schemas \emph{prior to} gold-label commitment, using only multi-annotator criterion judgments. The diagnostic separates two failure modes: unstable criteria with hard-to-operationalize boundaries, and systematic overlap that blurs the boundaries between mutually exclusive categories. Applied to persuasive value extraction in commercial documents, we find that disagreement is not diffuse: instability concentrates in a few criteria, while nearly half of covered sentences activate multiple categories. These signals align with where domain experts disagree, yielding an evidence-based audit for tightening guidelines, revising category structure, or reconsidering the annotation paradigm.

主观任务标注诊断可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。