arXiv:2604.28122cs.CVcs.LG2026-04

用球形隐变量提升视觉模型的几何保真度

Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces

论文配图:Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
图 1 · 摘自论文原文
  • 采用球形分布隐变量,强制保留方向与几何语义
  • 在高压缩率下,深度估计等任务性能显著优于传统高斯瓶颈
  • 适合需要物理一致性的3D场景建模研究者

现代视觉世界建模系统依赖大容量架构和大规模数据生成合理运动,但常无法保持底层3D几何结构或物理一致的相机动态。关键瓶颈不仅在于模型容量,更在于用于编码几何结构的隐表示。我们提出S²VAE,一种以几何为导向的隐学习框架,专注于压缩并表征场景的潜在3D状态,包括相机运动、深度和点级结构,而非仅建模外观。基于视觉几何基底变换器(VGGT)的表示,我们引入一种新型变分自编码器,使用幂球形隐变量的乘积,显式在瓶颈中强制超球面结构,以在强压缩下保留方向性和几何语义。在深度估计、相机位姿恢复和点云重建任务中,我们证明几何对齐的超球形隐变量始终优于传统高斯瓶颈,尤其在高压缩率下表现更优。结果凸显了隐空间几何作为物理基底视觉与世界模型的第一性设计原则。

原文摘要 · Abstract (English)

Modern visual world modeling systems increasingly rely on high-capacity architectures and large-scale data to produce plausible motion, yet they often fail to preserve underlying 3D geometry or physically consistent camera dynamics. A key limitation lies not only in model capacity, but in the latent representations used to encode geometric structure. We propose S$^2$VAE, a geometry-first latent learning framework that focuses on compressing and representing the latent 3D state of a scene, including camera motion, depth, and point-level structure, rather than modeling appearance alone. Building on representations from a Visual Geometry Grounded Transformer (VGGT), we introduce a novel type of variational autoencoder using a product of Power Spherical latent distributions, explicitly enforcing hyperspherical structure in the bottleneck to preserve directional and geometric semantics under strong compression. Across depth estimation, camera pose recovery, and point cloud reconstruction, we show that geometry-aligned hyperspherical latents consistently outperform conventional Gaussian bottlenecks, particularly in high-compression regimes. Our results highlight latent geometry as a first-class design choice for physically grounded visual and world models.

视觉建模隐变量几何感知自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。