构建首个开源三模态数据集,提升模型对手绘草图的理解能力
O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
- 基于图像-草图-指令三元组构建大规模训练数据集
- 在多项草图任务中超越现有模型,最高提升17.3%准确率
- 适合需要草图理解的AI绘画、交互设计等场景
尽管大型视觉语言模型(LVLM)在实际应用中日益普及,但其对抽象视觉输入的解析能力仍有限,尤其难以理解手绘草图——这种表达文字难以描述概念的直观方式。我们识别出主要瓶颈在于缺乏同时包含草图、真实图像和自然语言指令的大规模数据集。为此,本文提出两项关键贡献:(1) 构建一个大规模图像-草图-指令三元组数据集,用于预训练与指令微调;(2) 基于该数据集训练的O3SLM模型。在多个草图任务上的综合评估显示:(a) 物体定位,(b) 计数,(c) 图像检索(即SBIR与细粒度SBIR),(d) 视觉问答(VQA);结合现有三个草图数据集(QuickDraw!、Sketchy、Tu Berlin)及自建的SketchVCL数据集,O3SLM在各项任务中均达到当前最优性能,显著优于现有LVLM在草图理解与推理方面的表现。
原文摘要 · Abstract (English)
While Large Vision Language Models (LVLMs) are increasingly deployed in real-world applications, their ability to interpret abstract visual inputs remains limited. Specifically, they struggle to comprehend hand-drawn sketches, a modality that offers an intuitive means of expressing concepts that are difficult to describe textually. We identify the primary bottleneck as the absence of a large-scale dataset that jointly models sketches, photorealistic images, and corresponding natural language instructions. To address this, we present two key contributions: (1) a new, large-scale dataset of image-sketch-instruction triplets designed to facilitate both pretraining and instruction tuning, and (2) O3SLM, an LVLM trained on this dataset. Comprehensive evaluations on multiple sketch-based tasks: (a) object localization, (b) counting, (c) image retrieval i.e., (SBIR and fine-grained SBIR), and (d) visual question answering (VQA); while incorporating the three existing sketch datasets, namely QuickDraw!, Sketchy, and Tu Berlin, along with our generated SketchVCL dataset, show that O3SLM achieves state-of-the-art performance, substantially outperforming existing LVLMs in sketch comprehension and reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。