arXiv:2503.11096cs.CVcs.AI2025-03被引 2

用大模型辅助标注,人画框、AI填标签,提速省力。

Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation

  • 人只画框,大模型自动生成标签,分工明确
  • 在物体识别等任务中表现良好,具备跨任务泛化能力
  • 适合需要大规模标注的视觉项目,尤其缓解人工疲劳

传统图像标注依赖人工进行目标选择和标签分配,耗时且易因疲劳导致效率下降。本文提出一种新型人机协同框架,利用大模态模型(如GPT)的视觉理解能力辅助标注流程。人类标注者仅需通过边界框选择目标,而大模型自动完成相关标签生成。该框架显著降低人工的认知与时间负担,提升标注效率。通过分析不同标注任务下的系统表现,证明其可泛化至物体识别、场景描述及细粒度分类等任务。本方法展示了重构标注工作流的潜力,为计算机视觉中的大规模数据标注提供可扩展、高效的解决方案。最后讨论了将大模态模型融入标注流水线如何促进人机双向对齐,并应对信息过载带来的‘无限标注’困境,通过将部分任务转移给AI来缓解压力。

原文摘要 · Abstract (English)

Traditional image annotation tasks rely heavily on human effort for object selection and label assignment, making the process time-consuming and prone to decreased efficiency as annotators experience fatigue after extensive work. This paper introduces a novel framework that leverages the visual understanding capabilities of large multimodal models (LMMs), particularly GPT, to assist annotation workflows. In our proposed approach, human annotators focus on selecting objects via bounding boxes, while the LMM autonomously generates relevant labels. This human-AI collaborative framework enhances annotation efficiency by reducing the cognitive and time burden on human annotators. By analyzing the system's performance across various types of annotation tasks, we demonstrate its ability to generalize to tasks such as object recognition, scene description, and fine-grained categorization. Our proposed framework highlights the potential of this approach to redefine annotation workflows, offering a scalable and efficient solution for large-scale data labeling in computer vision. Finally, we discuss how integrating LMMs into the annotation pipeline can advance bidirectional human-AI alignment, as well as the challenges of alleviating the "endless annotation" burden in the face of information overload by shifting some of the work to AI.

人机协作图像标注大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。