用大模型当侦探,挑出可疑标注,让人工只改重点,省90%时间
ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
- 大模型自检标注错误,只让人工处理高风险样本
- 在多数数据集上性能损失低于2%,人力成本降低90%
- 适用于文本、视觉、多模态任务,提供可落地的标注指南
监督学习依赖高质量标注数据,但人工标注成本高、耗时长。现有方法尝试用大语言模型(LLM)进行标注,但质量仍不及人类。为此,我们提出批判性思维标注(ACT)流程:大模型不仅充当标注者,还作为评判者,主动识别潜在错误,使人工仅需审查最可疑的样本,大幅提高效率。主要贡献包括:(1) ACT 适用于自然语言处理(NLP)、计算机视觉(CV)及多模态理解等多种领域,依托多模态大语言模型(MLLMs)实现;(2) 通过实证研究总结7条提升标注质量并降低人力成本的关键洞察,并转化为易用指南;(3) 理论分析如何调整损失函数,使在ACT数据上训练的模型表现接近全人工标注数据训练的结果。实验表明,在多数基准数据集上性能差距小于2%,同时节省高达90%的人力成本。
原文摘要 · Abstract (English)
Supervised learning relies on high-quality labeled data, but obtaining such data through human annotation is both expensive and time-consuming. Recent work explores using large language models (LLMs) for annotation, but LLM-generated labels still fall short of human-level quality. To address this problem, we propose the Annotation with Critical Thinking (ACT) data pipeline, where LLMs serve not only as annotators but also as judges to critically identify potential errors. Human effort is then directed towards reviewing only the most "suspicious" cases, significantly improving the human annotation efficiency. Our major contributions are as follows: (1) ACT is applicable to a wide range of domains, including natural language processing (NLP), computer vision (CV), and multimodal understanding, by leveraging multimodal-LLMs (MLLMs). (2) Through empirical studies, we derive 7 insights on how to enhance annotation quality while efficiently reducing the human cost, and then translate these findings into user-friendly guidelines. (3) We theoretically analyze how to modify the loss function so that models trained on ACT data achieve similar performance to those trained on fully human-annotated data. Our experiments show that the performance gap can be reduced to less than 2% on most benchmark datasets while saving up to 90% of human costs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。