arXiv:2605.19066cs.CL2026-05中稿 · Deep Learning Inda…被引 3

低资源NLP评估面临标注人力稀缺与技术增速不匹配的矛盾。

The Annotation Scarcity Paradox in Low-Resource NLP Evaluation: A Decade of Acceleration and Emerging Constraints

  • 提出标注稀缺悖论,揭示评估能力滞后于模型扩展
  • 指出当前评估依赖廉价隐性劳动,威胁研究可信度
  • 倡导社区共建、数据主权的新型评估范式

过去十年,低资源自然语言处理(NLP)在跨语言迁移、多语言大模型和基准测试快速扩张的推动下迅猛发展。然而,这种表面进展掩盖了一个关键而被忽视的矛盾:评估日益复杂的生成系统所需的深层社会语言学专业知识严重不足、分布不均且结构性边缘化。本文对2014年以来低资源NLP评估进行批判性综述,梳理其三阶段演变:早期启发式乐观、自上而下基准扩增的幻象,以及当前生成瓶颈期。提出“标注稀缺悖论”概念,即技术扩容能力远超所需主权人力评估基础设施所形成的结构性摩擦。通过分析抽取式数据流水线、未获补偿的“幽灵工作”及语言数据激增现象,论证该悖论正威胁报告进展的真知性。综述新兴应对策略——包括数据增强、基于模型的评估、参与式整理,以及基于项目反应理论与主动学习的标注高效方法,并评估其公平性与有效性权衡。最后呼吁实践者行动:突破交易式数据提取,转向以认知治理、数据主权和共享所有权为基础的、嵌入社区的关系型评估范式。

原文摘要 · Abstract (English)

Over the past decade, low-resource natural language processing (NLP) has experienced explosive growth, propelled by cross-lingual transfer, massively multilingual models, and the rapid proliferation of benchmarks. Yet this apparent progress masks a critical, insufficiently examined tension: the deep sociolinguistic expertise required to evaluate increasingly complex generative systems is severely strained, inequitably distributed, and structurally marginalised. We present a critical narrative survey of low-resource NLP evaluation (2014-present), tracing its evolution across three phases: early heuristic optimism, the illusions of top-down benchmark scaling, and the current era of generative bottlenecks. We conceptualise the Annotation Scarcity Paradox, the structural friction arising when the technical capacity to scale models vastly outpaces the sovereign human infrastructure required to authentically evaluate them. By examining extractive data pipelines, undercompensated ``ghost work'', and language data flaring, we argue that this paradox threatens the epistemic validity of reported progress. We survey emerging responses -- including data augmentation, model-based evaluation, participatory curation, and annotation-efficient approaches via item response theory and active learning -- and assess their equity and validity trade-offs. We close with a practitioner call to action, arguing that overcoming this bottleneck requires a paradigm shift from transactional data extraction to relational, community-embedded evaluation rooted in epistemic governance, data sovereignty, and shared ownership.

低资源NLP评估困境数据主权社区共建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。