通过输入操作探索特征空间几何,发现线性映射可有效还原图像变换后的特征。
FeatMap: Understanding image manipulation in the feature space and its implications for feature space geometry

- 在输入空间施加多种变换,学习特征图间的映射关系。
- 全局模型表现最优,但共享线性模型也能接近效果,重建误差小。
- 揭示特征空间近似线性结构,适合研究模型内部机理的学者参考。
中间特征表示是深度神经网络表达能力和适应性的核心,但其几何结构仍不明确。本文通过在输入空间施加多种操作(包括几何、光度变换、局部掩码及生成式编辑模型的语义修改),评估从原始特征图到变换后特征图的映射学习可行性。设计了从线性到非线性、局部到全局的多种映射方式,评估映射的重建质量与语义内容。结果表明,所有变换均可学习有效映射;尽管全局(如Transformer)模型表现更优,但共享线性模型在单个特征向量上即可实现接近性能,即使面对复杂语义操作也仅有微小退化。分析不同特征层的映射,发现权重与偏置主导性及线性变换的有效秩具有层级差异。这些结果支持特征空间在第一近似下呈线性结构的假设。从更广视角看,生成式图像编辑模型为通过输入操控深入理解特征空间提供了新路径。
原文摘要 · Abstract (English)
Intermediate feature representations represent the backbone for the expressivity and adaptability of deep neural networks. However, their geometric structure remains poorly understood. In this submission, we provide indirect insights into this matter by applying a broad selection of manipulations in input space, ranging from geometric and photometric transformations to local masking and semantic manipulations using generative image editing models, and assess the feasibility of learning a mapping in the feature space, mapping from the original to the manipulated feature map. To this end, we devise different types of mappings, from linear to non-linear and local to global mappings and assess both the reconstruction quality of the mapping as well as the semantic content of the mapped representations. We demonstrate the feasibility of learning such mappings for all considered transformations. While global (transformer) models that operate on the full feature map often achieve best results, we show that the same can be achieved with a shared linear model operating on a single feature vector typically with very little degradation in reconstruction quality, even for highly non-trivial semantic manipulations. We analyze the corresponding mappings across different feature layers and characterize them according to dominance of weight vs. bias and the effective rank of the linear transformations. These results provide hints for the hypothesis that the feature space is to a first degree of approximation organized in linear structures. From a broader perspective, the study demonstrates that generative image editing models might open the door to a deeper understanding of the feature space through input manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。