早期内容质检比后期验证更省钱,能大幅降低错误成本。
Position: Early-Stage Quality Assurance in Annotation Pipelines Is More Cost-Effective Than Late-Stage Validation
- 提出标注流程三个质检时机:标注前、标注后、审校后。
- 实证表明早期发现错误可节省4至100倍成本。
- 呼吁学界公开质检时间点,推动平台支持时间参数配置。
本文主张机器学习领域应优先关注标注流程中的早期质量保障,而非当前普遍采用的后期验证。数据质量瓶颈正制约大模型发展,但现有研究几乎只聚焦验证方法,忽视验证时机。事实上,何时进行验证直接影响错误率与成本。借鉴软件工程中‘左移’原则,研究表明后期缺陷修复成本是前期的4至100倍(Boehm, 1981;Shull et al., 2002)。标注流程亦呈现类似规律:在标注前发现问题的成本仅为审校完成后的一小部分。本文提出三类质检触发点(T0预标注、T1标注后、T2审校后),将标注流程拆解为可操作的验证节点。通过参数化误差传播模型,明确时机对最终错误率与经济性的影响,使时机成为可度量的设计变量。对47篇近期论文的调查显示,仅有4%报告了验证时机,暴露显著缺口。忽略质检时机,可能导致在方法优化上投入过多,却忽视最具影响的结构性因素。实现该主张需三步:研究者应同时报告验证时机与方法;标注平台应将时机设为首要参数;社区应开展控制实验,直接测量各阶段检测率。
原文摘要 · Abstract (English)
This position paper argues that the machine learning community should prioritize early-stage quality assurance in annotation pipelines over the prevailing practice of late-stage validation. Data quality bottlenecks increasingly limit foundation model improvement, yet quality assurance research focuses almost exclusively on validation methods rather than validation timing. When validation occurs, not merely what methods are employed, fundamentally determines both error rates and annotation costs. This temporal neglect is puzzling given the well-established "shift-left" principle from software engineering, where empirical studies demonstrate 4--100x cost multipliers for defects detected in later stages (Boehm, 1981; Shull et al., 2002). Annotation pipelines exhibit analogous dynamics: errors caught before annotation begins cost a fraction of those discovered after review cycles complete. We propose a taxonomy of three QA trigger points, namely pre-annotation (T0), post-annotation (T1), and post-review (T2), that decompose annotation workflows into discrete validation opportunities. A parametric error-propagation model formalizes when timing affects final error rates versus only economics, making timing a measurable design variable rather than a configuration afterthought. A survey of 47 recent papers reveals that only 4% report when validation occurs, a striking gap given timing's demonstrated impact in adjacent fields. Without explicit attention to QA timing, the community risks optimizing validation methods while ignoring the structural variable that may matter most. Acting on this position requires three steps: researchers should report QA timing configurations alongside validation methods; annotation platforms should expose timing as a first-class parameter; and the community should run controlled experiments that measure stage-specific detection rates directly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。