arXiv:2607.15536cs.CVcs.LG2026-07

将颜色视为几何量,实现3D高斯点的旋转对称性统一建模。

E3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry Embedding

论文配图:E3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry Embedding
图 1 · 摘自论文原文
  • 把球谐系数映射为3×3矩阵,统一几何与外观的旋转变换
  • 在物体视觉和动作条件建模任务中,相机视角变化下性能更稳定
  • 无需张量积,适合追求对称性的3D生成与重建场景

3D高斯点阵(3DGS)通过显式几何(位置、协方差)与视图相关的外观(球谐系数)联合表征场景。然而,在这些原始数据上构建SE(3)等变架构存在根本性表示瓶颈:颜色被当作信号而非几何实体,导致在相机坐标系变化时几何与外观的对称性难以统一。虽然平移可通过相对坐标处理,但旋转对各属性作用不一致:μ→Rμ,Σ→RΣR^T,fℓ→Dℓ(R)fℓ。这种差异使严格等变难以实现,现有方法或丢弃或展平球谐系数,破坏对称性。本文基于表示论提出统一解法:当球谐阶数ℓ≤2时,外观与秩-2几何张量代数同构。我们证明,Wigner-D作用可精确重写为3×3矩阵的共轭作用。由此引入统一矩阵嵌入(Unified Matrix Embedding),将所有高斯属性映射至统一载体空间gl(3)。基于“颜色即几何”思想,提出E3DGS,一种无需克莱布什-戈登张量积的刚体(SE(3))等变架构。在物体视觉与动作条件高斯世界建模任务上的评估表明,该方法在相机帧变化下具备强鲁棒性与更高数据效率。

原文摘要 · Abstract (English)

3D Gaussian Splatting (3DGS) captures scenes by coupling explicit geometry (position, covariance) with view-dependent photometry (Spherical Harmonics). However, building $\mathrm{SE}(3)$-equivariant architectures on these primitives presents a fundamental representation bottleneck. Color has been treated as a signal rather than a geometric entity, making it nontrivial to unify symmetry across geometry and appearance as the camera frame changes. While translations are handled by relative coordinates, rotations act heterogeneously across attributes: $μ\mapsto Rμ$, $Σ\mapsto RΣR^\top$, and $f_\ell\mapsto D^\ell(R)f_\ell$. This mismatch complicates strict equivariance, leading existing methods to either discard or flatten SH coefficients, thereby breaking symmetry. We propose a unified solution rooted in representation theory: for SH degrees $\ell\le2$, photometry is algebraically isomorphic to a rank-2 geometric tensor. We prove that the Wigner-$D$ action on these SH coefficients can be exactly reformulated as the conjugation action on $3\times3$ matrices. Leveraging this, we introduce the Unified Matrix Embedding, a lifting that maps all Gaussian attributes into a unified carrier space, $\mathfrak{gl}(3)$. Building on the "Color-as-Geometry" formulation, we present E3DGS, a rigid-body ($\mathrm{SE}(3)$) equivariant architecture that processes 3D Gaussians without Clebsch-Gordan tensor products. Evaluations on object vision and action-conditioned Gaussian world modeling demonstrate that our unified approach yields strong robustness under camera-frame changes and improved data efficiency.

3D高斯等变网络球谐系数几何学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。