量化评估NLP实验可复现性,让结果更可信、可比较。
QRA++: Quantified Reproducibility Assessment for Common Types of Results in Natural Language Processing
- 提出三级粒度的连续可复现度评分方法
- 发现实验相似性越高,复现效果越好
- 适合关注模型可靠性与实验设计的研究者
NLP领域的复现研究虽已揭示该领域可复现性令人担忧,但每项研究基于隐含标准得出的定性结论难以比较和总结。本文提出QRA++,一种量化可复现性评估方法:(i)在三个粒度层级上生成连续值的可复现度评分;(ii)使用跨研究可比的复现度量;(iii)将复现预期建立在实验属性相似性的基础上。通过应用于三组可比实验,QRA++揭示可复现性显著依赖于实验属性相似性,同时也受系统类型和评估方法影响,为理解复现失败根源提供新视角。
原文摘要 · Abstract (English)
Reproduction studies reported in NLP provide individual data points which in combination indicate worryingly low levels of reproducibility in the field. Because each reproduction study reports quantitative conclusions based on its own, often not explicitly stated, criteria for reproduction success/failure, the conclusions drawn are hard to interpret, compare, and learn from. In this paper, we present QRA++, a quantitative approach to reproducibility assessment that (i) produces continuous-valued degree of reproducibility assessments at three levels of granularity; (ii) utilises reproducibility measures that are directly comparable across different studies; and (iii) grounds expectations about degree of reproducibility in degree of similarity between experiments. QRA++ enables more informative reproducibility assessments to be conducted, and conclusions to be drawn about what causes reproducibility to be better/poorer. We illustrate this by applying QRA++ to three example sets of comparable experiments, revealing clear evidence that degree of reproducibility depends on similarity of experiment properties, but also system type and evaluation method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。