arXiv:2410.09416cs.CVcs.AI2024-10中稿 · NeurIPS被引 7

视觉语言模型可替代人工标注,成本不足1%且准确率超89%。

Can Vision-Language Models Replace Human Annotators: A Case Study with CelebA Dataset

  • 用LLaVA-NeXT模型自动标注1000张人脸图,初步一致率达79.5%
  • 通过多数投票重标注分歧项,一致性提升至89.1%以上
  • 成本低于人工标注1%,适合大规模数据标注任务

本研究评估了视觉语言模型(VLMs)在图像数据标注中的能力,对比其在CelebA数据集上与人工标注在质量与成本效益方面的表现。使用最先进的LLaVA-NeXT模型对1000张CelebA图像进行标注,与原始人工标注的吻合度达79.5%。将存在分歧的样本进行重标注并采用多数投票后,AI标注的一致性提升至89.1%,客观标签甚至更高。成本分析显示,相较于传统人工方法,AI标注成本不足人工标注的1%。结果表明,VLMs在特定标注任务中具备可行性与成本优势,能显著降低财务负担与大规模人工标注带来的伦理问题。本研究使用的AI标注与重标注数据已公开于https://github.com/evev2024/EVEV2024_CelebA。

原文摘要 · Abstract (English)

This study evaluates the capability of Vision-Language Models (VLMs) in image data annotation by comparing their performance on the CelebA dataset in terms of quality and cost-effectiveness against manual annotation. Annotations from the state-of-the-art LLaVA-NeXT model on 1000 CelebA images are in 79.5% agreement with the original human annotations. Incorporating re-annotations of disagreed cases into a majority vote boosts AI annotation consistency to 89.1% and even higher for more objective labels. Cost assessments demonstrate that AI annotation significantly reduces expenditures compared to traditional manual methods -- representing less than 1% of the costs for manual annotation in the CelebA dataset. These findings support the potential of VLMs as a viable, cost-effective alternative for specific annotation tasks, reducing both financial burden and ethical concerns associated with large-scale manual data annotation. The AI annotations and re-annotations utilized in this study are available on https://github.com/evev2024/EVEV2024_CelebA.

视觉语言模型数据标注成本优化自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。