arXiv:2505.22627cs.CLcs.CV2025-05EMNLP被引 1

通过对话式顺序标注,提升图像描述的效率与全面性。

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions

  • 采用逐轮对话方式,后继标注者仅补充遗漏信息,减少重复工作。
  • 实验表明标注速度达0.42单位/秒,优于传统并行方法的0.30单位/秒。
  • 适合需要高质量密集图像描述的研究者快速构建数据集。

尽管密集标注的图像描述能显著促进视觉-语言对齐的学习,但系统优化人工标注效率的方法仍鲜有研究。本文提出链式对话标注(Chain-of-Talkers, CoTalk),一种人机协同框架,在固定预算(如总标注时间)下最大化标注样本数并提升内容完整性。该框架基于两大核心洞察:其一,序列化标注相比传统并行方式可减少冗余,后续标注者只需补充前序未覆盖的视觉信息;其二,人类阅读文本并口述输出的效率远高于打字,多模态接口实现更高吞吐。我们从内在评估和外在评估两方面验证框架效果:内在评估通过将详细描述解析为物体-属性树,分析语义单元的有效连接;外在评估则考察标注结果在视觉-语言对齐任务中的实际性能。八名参与者实验显示,CoTalk的标注速度为0.42单位/秒,显著高于并行方法的0.30单位/秒;检索准确率也从40.52%提升至41.13%。

原文摘要 · Abstract (English)

While densely annotated image captions significantly facilitate the learning of robust vision-language alignment, methodologies for systematically optimizing human annotation efforts remain underexplored. We introduce Chain-of-Talkers (CoTalk), an AI-in-the-loop methodology designed to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints (e.g., total human annotation time). The framework is built upon two key insights. First, sequential annotation reduces redundant workload compared to conventional parallel annotation, as subsequent annotators only need to annotate the ``residual'' -- the missing visual information that previous annotations have not covered. Second, humans process textual input faster by reading while outputting annotations with much higher throughput via talking; thus a multimodal interface enables optimized efficiency. We evaluate our framework from two aspects: intrinsic evaluations that assess the comprehensiveness of semantic units, obtained by parsing detailed captions into object-attribute trees and analyzing their effective connections; extrinsic evaluation measures the practical usage of the annotated captions in facilitating vision-language alignment. Experiments with eight participants show our Chain-of-Talkers (CoTalk) improves annotation speed (0.42 vs. 0.30 units/sec) and retrieval performance (41.13% vs. 40.52%) over the parallel method.

图像描述人机协同高效标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。