用艺术描述做训练上下文,提升艺术图像定位能力
Context-Infused Visual Grounding for Art
- 在训练中引入艺术文本描述作为上下文,增强模型对艺术图像的理解
- 在两个艺术数据集上实现新最优的物体检测性能
- 构建了首个乌贼绘浮世绘图文定位数据集Ukiyo-eVG
许多艺术作品集合包含丰富的文本属性,提供对作品的详细描述。视觉定位可将这些描述中的主体在图像中精确定位,但现有方法基于自然图像训练,在艺术图像上泛化能力差。本文提出CIGAr(Context-Infused GroundingDINO for Art),在训练中利用艺术作品描述作为上下文,实现艺术图像上的精准定位。同时,构建新数据集Ukiyo-eVG,包含人工标注的短语-图像对应关系,并在两个艺术数据集上达到新的最优物体检测性能。
原文摘要 · Abstract (English)
Many artwork collections contain textual attributes that provide rich and contextualised descriptions of artworks. Visual grounding offers the potential for localising subjects within these descriptions on images, however, existing approaches are trained on natural images and generalise poorly to art. In this paper, we present CIGAr (Context-Infused GroundingDINO for Art), a visual grounding approach which utilises the artwork descriptions during training as context, thereby enabling visual grounding on art. In addition, we present a new dataset, Ukiyo-eVG, with manually annotated phrase-grounding annotations, and we set a new state-of-the-art for object detection on two artwork datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。