用傅里叶形状探索神经网络的几何感知能力
Learning Fourier shapes to probe the geometric world of deep neural networks
- 用傅里叶级数参数化任意形状,端到端可微优化
- 几何形状可高置信度分类,且精准定位模型关注区域
- 揭示新对抗范式,适合研究模型可解释性与鲁棒性
尽管形状和纹理都是视觉识别的基础,但深度神经网络(DNN)研究长期侧重于纹理,对几何理解不足。本文首次证明:优化后的形状可作为强语义载体,仅凭几何信息即可生成高置信度分类;它们是高保真可解释性工具,能精确分离模型显著区域;并构成一种新型通用对抗范式,可欺骗下游视觉任务。方法基于端到端可微框架,结合强大的傅里叶级数参数化任意形状、基于绕数的像素网格映射,以及信号能量约束以提升优化效率并保证物理合理性。本工作为探测DNN几何世界提供通用工具,开辟机器感知理解的新前沿。
原文摘要 · Abstract (English)
While both shape and texture are fundamental to visual recognition, research on deep neural networks (DNNs) has predominantly focused on the latter, leaving their geometric understanding poorly probed. Here, we show: first, that optimized shapes can act as potent semantic carriers, generating high-confidence classifications from inputs defined purely by their geometry; second, that they are high-fidelity interpretability tools that precisely isolate a model's salient regions; and third, that they constitute a new, generalizable adversarial paradigm capable of deceiving downstream visual tasks. This is achieved through an end-to-end differentiable framework that unifies a powerful Fourier series to parameterize arbitrary shapes, a winding number-based mapping to translate them into the pixel grid required by DNNs, and signal energy constraints that enhance optimization efficiency while ensuring physically plausible shapes. Our work provides a versatile framework for probing the geometric world of DNNs and opens new frontiers for challenging and understanding machine perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。