形式语义结构只能解释人类标注差异中极小部分,对分歧程度和内容影响有限。
How Much Human Label Variation Does Formal Semantic Structure Explain?: Group-Level Effects and Item-Level Ceilings in NLI
- 用规则系统与单调性标签器分析3113个样本的语义结构
- 语义结构仅解释3.3%~3.6%的标注差异方差,判别力弱
- 无法预测高分歧项,且不影响分歧内容类型
人类在自然语言推理中的标注差异日益被视为有效信号,但其受形式语义结构影响的程度尚未直接测量。本研究基于ChaosNLI的3,113个SNLI与MNLI样本,采用经MED验证(编辑层面一致性0.883,句子摘要层面0.807)的规则化算子与单调性标签器,通过三个预注册分析模块及负结果完整报告进行分析。发现三重边界:第一,群体层面边界显示,非纯上向单调的假设具有显著更高标注熵(Cliff's delta = -0.284),该效应在秩检验中稳健,尽管长度控制的回归形式被敏感性检验削弱;第二,个体层面天花板显示,相同形式特征仅解释3.3%至3.6%的熵方差,中位数分组AUC为0.606,不足以识别高分歧项目;第三,组合不变性表明,在边界两侧,三个高统计功效的预注册对比(基于错误率与解释类型份额,即VariErr、LiTEx)均未发现显著差异。在该样本范围内,形式语义结构仅小幅改变标注者分歧程度,且不明显改变分歧内容。所有结论均基于低原始一致性的ChaosNLI-S/M数据集,并受此范围约束。所有分析均已预注册于版本控制的研究日志,包含一次修正的解释规则,论文已公开审计轨迹。
原文摘要 · Abstract (English)
Human label variation in natural language inference is increasingly treated as signal rather than noise, but how much of it formal semantic structure explains has not been measured directly. We measure it on the 3,113 SNLI and MNLI items of ChaosNLI, using a rule-based operator and monotonicity tagger validated against MED (0.883 agreement at the edit site, 0.807 on the sentence-level summary our analyses consume), three preregistered analysis blocks, and full reporting of negative results. Three bounds emerge. First, a group-level boundary: hypotheses that are not purely upward monotone show reliably higher label entropy (Cliff's delta = -0.284), and rank-based tests defend the effect against operator-presence and length reductions, though a bounded-outcome sensitivity check weakens the regression form of the length defense. Second, an item-level ceiling: the same formal profiles explain only 3.3 to 3.6 percent of entropy variance and reach a median-split AUC of 0.606, too weak to identify high-disagreement items. Third, composition invariance: across the boundary, three high-powered preregistered contrasts on validated error shares and explanation-type shares (VariErr, LiTEx) all return null results. In this sample, formal semantic structure shifts how much annotators disagree by a small amount and does not detectably change what they disagree about. ChaosNLI-S/M consists of items selected for low original agreement, and every claim is conditioned on that scope. All analyses were preregistered in a version-controlled research log, whose audit trail, including one corrected interpretation rule, the paper discloses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。