让神经网络学会处理局部对称性,提升3D视觉识别效果
Learning from Frustration: Torsor CNNs on Graphs
- 用边上的群变换建模局部坐标系变化,实现局部等变
- 引入挫折损失函数,自动优化局部等变表示
- 适用于多视角3D识别,无需全局坐标系
大多数等变神经网络依赖单一全局对称性,限制了其在局部对称性场景中的应用。本文提出Torsor CNN框架,通过边上的势能(群值变换)编码图上的局部对称性。该几何构造与经典群同步问题本质等价,由此得出:(1) 可证明局部坐标系变化下保持等变的Torsor卷积层;(2) 挫折损失——一种独立的几何正则化项,可添加至任意神经网络训练目标中,以促进局部等变表示。该框架统一并推广了多个架构(包括经典CNN和流形上的规范CNN),可在任意图上运行,无需全局坐标系或光滑流形结构。我们建立了该框架的数学基础,并在多视角3D识别任务中验证其有效性,其中相对相机位姿自然定义所需的边势能。
原文摘要 · Abstract (English)
Most equivariant neural networks rely on a single global symmetry, limiting their use in domains where symmetries are instead local. We introduce Torsor CNNs, a framework for learning on graphs with local symmetries encoded as edge potentials -- group-valued transformations between neighboring coordinate frames. We establish that this geometric construction is fundamentally equivalent to the classical group synchronization problem, yielding: (1) a Torsor Convolutional Layer that is provably equivariant to local changes in coordinate frames, and (2) the frustration loss -- a standalone geometric regularizer that encourages locally equivariant representations when added to any NN's training objective. The Torsor CNN framework unifies and generalizes several architectures -- including classical CNNs and Gauge CNNs on manifolds -- by operating on arbitrary graphs without requiring a global coordinate system or smooth manifold structure. We establish the mathematical foundations of this framework and demonstrate its applicability to multi-view 3D recognition, where relative camera poses naturally define the required edge potentials.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。