arXiv:2508.06696cs.CV2025-08

用线条图先学结构,让模型更省力、更通用、更像人。

Learning More by Seeing Less: Structure First Learning for Efficient, Transferable, and Human-Aligned Vision

  • 用线稿作为初始训练数据,引导模型关注形状而非颜色细节
  • 在分类、检测、分割任务中均提升数据效率,且表征维度更低
  • 生成的模型更易压缩,适合部署到轻量级设备上

尽管计算机视觉取得显著进展,现代识别系统仍严重依赖丰富冗余的视觉输入。相比之下,人类能轻松理解线条图等稀疏表示,表明结构而非外观才是高效视觉理解的基础。本文提出一种结构优先学习范式,以线条图为初始训练模态,诱导更紧凑、泛化性更强的视觉表征。实验表明,此类模型具有更强的形状偏好、更聚焦的注意力,且在分类、检测和分割任务中数据效率更高。值得注意的是,这些模型内在维度更低,需更少主成分即可捕捉表征方差,与人类大脑低维高效表征现象一致。此外,结构优先学习产生更可压缩的表征,有利于向轻量学生模型蒸馏。从线稿教师蒸馏的学生模型始终优于从彩色监督教师蒸馏的模型,凸显结构紧凑知识的优势。结果支持结构优先学习能促进效率、泛化性和类人归纳偏置,为构建更鲁棒、自适应的视觉系统提供简单而有力的方法。

原文摘要 · Abstract (English)

Despite remarkable progress in computer vision, modern recognition systems remain fundamentally limited by their dependence on rich, redundant visual inputs. In contrast, humans can effortlessly understand sparse, minimal representations like line drawings, suggesting that structure, rather than appearance, underlies efficient visual understanding. In this work, we propose a novel structure-first learning paradigm that uses line drawings as an initial training modality to induce more compact and generalizable visual representations. We demonstrate that models trained with this approach develop a stronger shape bias, more focused attention, and greater data efficiency across classification, detection, and segmentation tasks. Notably, these models also exhibit lower intrinsic dimensionality, requiring significantly fewer principal components to capture representational variance, which mirrors observations of low-dimensional, efficient representations in the human brain. Beyond performance improvements, structure-first learning produces more compressible representations, enabling better distillation into lightweight student models. Students distilled from teachers trained on line drawings consistently outperform those trained from color-supervised teachers, highlighting the benefits of structurally compact knowledge. Together, our results support the view that structure-first visual learning fosters efficiency, generalization, and human-aligned inductive biases, offering a simple yet powerful strategy for building more robust and adaptable vision systems.

结构学习数据效率模型压缩类人认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。