提出新数据集,让模型更好识别因果关系的反例。
Investigating Counterclaims in Causality Extraction from Text
- 构建包含因果/反因果/非因果三类语句的标注体系
- 模型在反例上误判率降低至原模型的1/10以下
- 适合研究因果推理、批判性文本分析的学者使用
许多因果主张(如“糖导致多动”)存在争议或已被推翻,但现有因果抽取研究几乎完全忽略反因果表述。为填补这一空白,我们系统梳理了因果抽取文献,整理出丰富的反因果语言表达形式,并制定明确的标注规范。基于此规范,我们构建了一个新数据集,包含1028条因果陈述、952条反因果陈述和1435条非因果陈述,达成较高的标注一致性(Cohen's κ=0.74)。实验表明,仅用因果陈述训练的先进模型对反因果陈述的误判率是使用本数据集训练模型的10倍以上。
原文摘要 · Abstract (English)
Many causal claims, such as "sugar causes hyperactivity," are disputed or outdated. Yet research on causality extraction from text has almost entirely neglected counterclaims of causation. To close this gap, we conduct a thorough literature review of causality extraction, compile an extensive inventory of linguistic realizations of countercausal claims, and develop rigorous annotation guidelines that explicitly incorporate countercausal language. We also highlight how counterclaims of causation are an integral part of causal reasoning. Based on our guidelines, we construct a new dataset comprising 1028 causal claims, 952 counterclaims, and 1435 uncausal statements, achieving substantial inter-annotator agreement (Cohen's $κ= 0.74$). In our experiments, state-of-the-art models trained solely on causal claims misclassify counterclaims more than 10 times as often as models trained on our dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。