arXiv:2511.12452cs.CVcs.CL2025-11被引 1

用语音生成密集图文描述,大幅提升图像与3D场景标注效率。

DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions

  • 通过语音实时关联视觉区域,实现高效细粒度标注
  • 构建含7460个3D物体的多语言数据集,支持跨文化与3D理解
  • 训练模型在多语言、文化对齐和3D空间能力上提升超50%

随着多模态大模型广泛应用,高质量任务导向训练数据需求迫切。现有数据集多依赖网络爬取或手动打字的稀疏标注,难以覆盖图像全部视觉内容。密集标注价值高但稀缺。传统文本标注方式表达受限、速度慢,尤其在跨文化图像与3D资产标注中表现不佳。本文提出DenseAnnotate,一个音频驱动的在线标注平台,允许标注员口述观察并同步标记对应视觉区域或3D部分。平台集成语音转文字与注意力区域定位。我们在两个领域开展案例研究,涉及超1000名标注员:跨文化图像与3D场景。最终构建一个包含3531张图像、898个3D场景、7460个3D对象的多模态数据集,覆盖20种语言,含8746条图像描述、2000条场景描述、19000条物体描述。基于该数据训练的模型,在多语言任务上提升5%,文化对齐提升47%,3D空间理解提升54%。结果表明DenseAnnotate为未来视觉-语言研究提供可行方案,可推广至多种任务与数据类型。

原文摘要 · Abstract (English)

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on sparse annotations mined from the Internet or entered via manual typing that capture only a fraction of an image's visual content. Dense annotations are more valuable but remain scarce. Traditional text-based annotation pipelines are poorly suited for creating dense annotations: typing limits expressiveness, slows annotation speed, and underrepresents nuanced visual features, especially in specialized areas such as multicultural imagery and 3D asset annotation. In this paper, we present DenseAnnotate, an audio-driven online annotation platform that enables efficient creation of dense, fine-grained annotations for images and 3D assets. Annotators narrate observations aloud while synchronously linking spoken phrases to image regions or 3D scene parts. Our platform incorporates speech-to-text transcription and region-of-attention marking. To demonstrate the effectiveness of DenseAnnotate, we conducted case studies involving over 1,000 annotators across two domains: culturally diverse images and 3D scenes. We curate a human-annotated multi-modal dataset of 3,531 images, 898 3D scenes, and 7,460 3D objects, with audio-aligned dense annotations in 20 languages, including 8,746 image captions, 2,000 scene captions, and 19,000 object captions. Models trained on this dataset exhibit improvements of 5% in multilingual, 47% in cultural alignment, and 54% in 3D spatial capabilities. Our results show that our platform offers a feasible approach for future vision-language research and can be applied to various tasks and diverse types of data.

多模态标注语音输入3D理解数据构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。