arXiv:2505.23763cs.CV2025-05CVPR被引 3

为手绘图设计高效网络,实现99%算力压缩仍保精度。

Sketch Down the FLOPs: Towards Efficient Networks for Human Sketch

  • 提出专用于手绘图的轻量化组件,可即插即用适配现有模型。
  • 在手绘检索任务中,算力降至0.254G(原40.18G),降幅99.37%。
  • 适合需低延迟部署的手绘识别场景,如智能绘画工具。

随着手绘研究日趋成熟,其大规模商业化应用已近在眼前。尽管图像领域的轻量模型已很成熟,但针对手绘数据的高效推理仍无研究。本文首次证明,现有适用于图像的轻量模型无法有效处理手绘数据。为此,我们提出两个专为手绘设计的模块:一是跨模态知识蒸馏网络,将图像轻量模型迁移至手绘领域,使算力和参数分别减少97.96%和84.89%;二是基于强化学习的画布选择器,利用手绘抽象特性动态调整抽象层级,进一步降低算力达三分之二。最终整体算力从40.18G降至0.254G(降幅99.37%),精度保持在33.03%(原为32.77%),实现比最优图像模型更低算力的手绘高效网络。

原文摘要 · Abstract (English)

As sketch research has collectively matured over time, its adaptation for at-mass commercialisation emerges on the immediate horizon. Despite an already mature research endeavour for photos, there is no research on the efficient inference specifically designed for sketch data. In this paper, we first demonstrate existing state-of-the-art efficient light-weight models designed for photos do not work on sketches. We then propose two sketch-specific components which work in a plug-n-play manner on any photo efficient network to adapt them to work on sketch data. We specifically chose fine-grained sketch-based image retrieval (FG-SBIR) as a demonstrator as the most recognised sketch problem with immediate commercial value. Technically speaking, we first propose a cross-modal knowledge distillation network to transfer existing photo efficient networks to be compatible with sketch, which brings down number of FLOPs and model parameters by 97.96% percent and 84.89% respectively. We then exploit the abstract trait of sketch to introduce a RL-based canvas selector that dynamically adjusts to the abstraction level which further cuts down number of FLOPs by two thirds. The end result is an overall reduction of 99.37% of FLOPs (from 40.18G to 0.254G) when compared with a full network, while retaining the accuracy (33.03% vs 32.77%) -- finally making an efficient network for the sparse sketch data that exhibit even fewer FLOPs than the best photo counterpart.

手绘识别轻量化模型知识蒸馏算力优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。