arXiv:2606.24734cs.CLcs.AI2026-06

将复杂标注任务拆解,降低整体认知负担,提升效率。

Task Decomposition for Efficient Annotation

论文配图:Task Decomposition for Efficient Annotation
图 1 · 摘自论文原文
  • 按中心实体分解标注任务,减少认知负荷
  • 拆解后可降低标注总成本,提升效率
  • 适合多角色协作的标注项目,如人机协同

高质量结构化标注在大规模语料上成本高昂。人工标注耗时,模型生成虽便宜但需昂贵验证和大量监督以保证质量。传统流程由单个标注者完成整例标注,但结构化任务复杂,各环节带来不同认知负担。现代标注项目可融合多样标注者(含模型与不同专业背景的人类)。如何在该背景下重新设计任务、合理分配工作仍不明确。本文提出将标注任务分解为子任务,以降低整体认知负荷。受中心化理论启发,构建基于有效标注空间自由度的推理负荷模型,发现识别中心实体(即子任务锚点)可压缩输出空间复杂度。通过隔离并优先推进中心识别的分解方式,显著降低总负荷。提供具体分解指南,并以实例展示成本效益提升。最后提出在固定预算下最优分配子任务给异构标注者的流程。

原文摘要 · Abstract (English)

High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is laborious, and model-based annotation, although cheaper to generate, requires expensive validation and potentially significant supervision to ensure that the annotation quality is strong enough to be useful downstream. In traditional annotation workflows, annotation of each complete example is performed end-to-end by a single annotator. However, structured annotation is complex, and each aspect of the task represents a unique challenge with an associated inferential load for a given annotator. Modern annotation projects can incorporate heterogeneous groups of annotators, including both models and human annotators with varying domain and linguistic expertise. It remains unclear, however, how to redesign annotation tasks in this setting, where efforts are discriminately allocated across heterogeneous annotators with respect to distinct annotation challenges. We propose to decompose annotation tasks into sub-tasks in order to reduce the aggregate inferential load of annotation projects. Inspired by the notion of centers from centering theory, we introduce a formal model of inferential load based on the degrees of freedom in the space of valid annotations. Using this model, we show that identifying these centers (i.e. salient anchor entities realized by annotation sub-tasks) constrains the output space complexity, and decompositions which isolate and advance center identification reduce the aggregate inferential load. We provide guidelines for decomposing complex structured annotation tasks, supported by examples demonstrating improved cost-efficiency from our prior work. Finally, we present a procedure for allocating sub-tasks across annotators to maximize quality under a fixed budget.

标注优化任务分解人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。