arXiv:2602.02564cs.LGcs.MA2026-02

用智能代理自动标注数据,提升准确率并评估标注者质量

Label Curation Using Agentic AI

  • 多智能体协同生成与验证标签,无需真实标签即可工作
  • 在低质标注者场景下,准确率最高提升50%
  • 可无监督评估标注者可靠性,适合大规模数据清洗

数据标注对监督学习至关重要,但随着数据规模和模态增长,生成准确、无偏且可扩展的标签仍具挑战。传统人工标注流程成本高、速度慢且易受标注者差异影响,推动了对可靠性感知的自动化标注需求。我们提出AURA(Agentic AI for Unified Reliability Modeling and Annotation Aggregation),一种用于大规模多模态数据标注的智能代理框架。AURA通过多个AI智能体协同生成与验证标签,无需真实标签。其核心是基于期望最大化(EM)算法的统计模型,联合推断隐藏的真实标签与标注者可靠性,利用混淆矩阵融合冲突标注并聚合噪声预测。在四个基准数据集上的实验表明,AURA相比基线最高提升5.8%准确率;在标注者质量较差的场景下,性能提升可达50%。此外,AURA能准确估计标注者可靠性,实现无需预验证的标注质量评估。

原文摘要 · Abstract (English)

Data annotation is essential for supervised learning, yet producing accurate, unbiased, and scalable labels remains challenging as datasets grow in size and modality. Traditional human-centric pipelines are costly, slow, and prone to annotator variability, motivating reliability-aware automated annotation. We present AURA (Agentic AI for Unified Reliability Modeling and Annotation Aggregation), an agentic AI framework for large-scale, multi-modal data annotation. AURA coordinates multiple AI agents to generate and validate labels without requiring ground truth. At its core, AURA adapts a classical probabilistic model that jointly infers latent true labels and annotator reliability via confusion matrices, using Expectation-Maximization to reconcile conflicting annotations and aggregate noisy predictions. Across the four benchmark datasets evaluated, AURA achieves accuracy improvements of up to 5.8% over baseline. In more challenging settings with poor quality annotators, the improvement is up to 50% over baseline. AURA also accurately estimates the reliability of annotators, allowing assessment of annotator quality even without any pre-validation steps.

数据标注智能代理可靠性建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。