多尺度协作特征提取,提升草图识别与生成精度
ViSketch-GPT: Collaborative Multi-Scale Feature Extraction for Sketch Recognition and Generation
- 采用多尺度上下文提取,融合不同层级特征协同工作
- 在QuickDraw数据集上分类与生成性能均超越现有方法
- 适合需要精准草图理解的视觉任务应用
理解人类草图的本质具有挑战性,因其创作方式差异大。准确识别复杂结构模式可同时提升草图识别精度与生成质量。本文提出ViSketch-GPT,一种基于多尺度上下文提取的新算法。该模型在多个尺度捕获精细细节,并通过类集成机制融合特征,使提取的特征协同增强关键细节的识别与生成。在QuickDraw数据集上的大量实验验证了其有效性,模型在分类与生成任务中均建立新基准,显著优于现有方法,准确率与生成草图保真度均有明显提升。该算法为理解复杂结构提供稳健框架,通过协作特征提取增强对草图等结构的理解,适用于计算机视觉与机器学习中的多种应用场景。
原文摘要 · Abstract (English)
Understanding the nature of human sketches is challenging because of the wide variation in how they are created. Recognizing complex structural patterns improves both the accuracy in recognizing sketches and the fidelity of the generated sketches. In this work, we introduce ViSketch-GPT, a novel algorithm designed to address these challenges through a multi-scale context extraction approach. The model captures intricate details at multiple scales and combines them using an ensemble-like mechanism, where the extracted features work collaboratively to enhance the recognition and generation of key details crucial for classification and generation tasks. The effectiveness of ViSketch-GPT is validated through extensive experiments on the QuickDraw dataset. Our model establishes a new benchmark, significantly outperforming existing methods in both classification and generation tasks, with substantial improvements in accuracy and the fidelity of generated sketches. The proposed algorithm offers a robust framework for understanding complex structures by extracting features that collaborate to recognize intricate details, enhancing the understanding of structures like sketches and making it a versatile tool for various applications in computer vision and machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。