arXiv:2604.19943cs.CL2026-04被引 1

健康素养标注中分歧有结构,不能简单聚合。

Structured Disagreement in Health-Literacy Annotation: Epistemic Stability, Conceptual Difficulty, and Agreement-Stratified Inference

论文配图:Structured Disagreement in Health-Literacy Annotation: Epistemic Stability, Conceptual Difficulty, and Agreement-Stratified Inference
图 1 · 摘自论文原文
  • 用比例评分法分析6323份新冠回应,捕捉标注差异分布
  • 任务本身概念难度比标注者差异影响更大,分歧是结构性的
  • 不同一致程度下社会差异方向可能相反,需分层分析

在自然语言处理中,标注常假设每条数据存在单一真实标签,通过聚合解决分歧。本文基于厄瓜多尔和秘鲁收集的6,323份开放性新冠回答,开展大规模分级健康素养标注分析。每条回答由多位标注者独立评分,采用比例正确性分数反映其与公共卫生规范的契合度,从而分析完整判断分布而非聚合标签。方差分解显示,问题层面的概念难度解释的变异远大于标注者身份,表明分歧主要源于任务本身而非个体差异。分一致性层次分析发现,国家、教育水平、城乡差异等社会科学效应在不同一致水平下大小不同,甚至方向反转。结果表明,分级健康素养评估兼具认知稳定与不稳定成分,简单聚合会掩盖重要推断差异。因此,强观点论建模不仅是概念合理,更是实现有效推断的统计必要。

原文摘要 · Abstract (English)

Annotation pipelines in Natural Language Processing (NLP) commonly assume a single latent ground truth per instance and resolve disagreement through label aggregation. Perspectivist approaches challenge this view by treating disagreement as potentially informative rather than erroneous. We present a large-scale analysis of graded health-literacy annotations from 6,323 open-ended COVID-19 responses collected in Ecuador and Peru. Each response was independently labeled by multiple annotators using proportional correctness scores, reflecting the degree to which responses align with normative public-health guidelines, allowing us to analyze the full distribution of judgments rather than aggregated labels. Variance decomposition shows that question-level conceptual difficulty accounts for substantially more variance than annotator identity, indicating that disagreement is structured by the task itself rather than driven by individual raters. Agreement-stratified analyses further reveal that key social-scientific effects, including country, education, and urban-rural differences, vary in magnitude and in some cases reverse direction across levels of inter-annotator agreement. These findings suggest that graded health-literacy evaluation contains both epistemically stable and unstable components, and that aggregating across them can obscure important inferential differences. We therefore argue that strong perspectivist modeling is not only conceptually justified but statistically necessary for valid inference in graded interpretive tasks.

健康素养标注分歧分级评分社会差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。