arXiv:2510.14536cs.CV2025-10中稿 · The 36th British M…

将图像分解为可解释的视觉特征,提升生成与编辑效果

Exploring Image Representation with Decoupled Classical Visual Descriptors

  • 将图像显式拆分为边缘、颜色等经典特征
  • 通过重建预训练学习各特征本质,保持可解释性
  • 适合需要可控生成的视觉任务,如图像编辑

探索高效图像表示是计算机视觉中的长期挑战。尽管深度学习在图像理解任务中取得显著进展,但其内部表征往往不透明,难以解释视觉信息如何被处理。相比之下,传统视觉描述符(如边缘、颜色和强度分布)长期以来是图像分析的基础,且对人类直观易懂。受此启发,本文提出核心问题:现代学习能否从这些经典线索中获益?我们以VisualSplit框架作答,该框架显式将图像分解为解耦的经典描述符,将其视为视觉知识的独立互补成分。通过基于重建的预训练方案,VisualSplit学习捕捉每个描述符的本质,同时保留可解释性。通过显式分解视觉属性,该方法天然支持在图像生成与编辑等高级视觉任务中实现有效属性控制,超越传统分类与分割任务,表明这种新学习范式在视觉理解中的有效性。

原文摘要 · Abstract (English)

Exploring and understanding efficient image representations is a long-standing challenge in computer vision. While deep learning has achieved remarkable progress across image understanding tasks, its internal representations are often opaque, making it difficult to interpret how visual information is processed. In contrast, classical visual descriptors (e.g. edge, colour, and intensity distribution) have long been fundamental to image analysis and remain intuitively understandable to humans. Motivated by this gap, we ask a central question: Can modern learning benefit from these classical cues? In this paper, we answer it with VisualSplit, a framework that explicitly decomposes images into decoupled classical descriptors, treating each as an independent but complementary component of visual knowledge. Through a reconstruction-driven pre-training scheme, VisualSplit learns to capture the essence of each visual descriptor while preserving their interpretability. By explicitly decomposing visual attributes, our method inherently facilitates effective attribute control in various advanced visual tasks, including image generation and editing, extending beyond conventional classification and segmentation, suggesting the effectiveness of this new learning approach for visual understanding. Project page: https://chenyuanqu.com/VisualSplit/.

图像表示可解释性视觉生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。