解决多视频实例标签不一致问题,实现稳定4D高斯点云动态建模。
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
- 引入可微排列潜变量,通过Sinkhorn层对齐不同视频的实例标签。
- 在Panoptic Studio上提升PSNR至28.36,实例mIoU达0.9129。
- 适合需要长期轨迹追踪与实例分离的动态场景重建任务。
我们提出Inst4DGS,一种具有长时序每高斯轨迹的实例分解4D高斯点云渲染方法。尽管动态4DGS发展迅速,但实例分解4DGS仍研究不足,主要因独立分割的多视角视频间实例标签不一致。为此,我们引入每视频标签排列潜变量,通过可微的Sinkhorn层学习跨视频实例匹配,实现一致身份下的直接多视角监督。该显式标签对齐使决策边界更清晰,且无身份漂移。为进一步提升效率,我们设计实例分解运动骨架,为每个物体提供低维运动基底以优化长时序轨迹。在Panoptic Studio和Neural3DV数据集上的实验表明,Inst4DGS同时支持跟踪与实例分解,渲染与分割质量均达领先水平。在Panoptic Studio上,其PSNR从26.10提升至28.36,实例mIoU从0.6310增至0.9129,超越最强基线。
原文摘要 · Abstract (English)
We present Inst4DGS, an instance-decomposed 4D Gaussian Splatting (4DGS) approach with long-horizon per-Gaussian trajectories. While dynamic 4DGS has advanced rapidly, instance-decomposed 4DGS remains underexplored, largely due to the difficulty of associating inconsistent instance labels across independently segmented multi-view videos. We address this challenge by introducing per-video label-permutation latents that learn cross-video instance matches through a differentiable Sinkhorn layer, enabling direct multi-view supervision with consistent identity preservation. This explicit label alignment yields sharp decision boundaries and temporally stable identities without identity drift. To further improve efficiency, we propose instance-decomposed motion scaffolds that provide low-dimensional motion bases per object for long-horizon trajectory optimization. Experiments on Panoptic Studio and Neural3DV show that Inst4DGS jointly supports tracking and instance decomposition while achieving state-of-the-art rendering and segmentation quality. On the Panoptic Studio dataset, Inst4DGS improves PSNR from 26.10 to 28.36, and instance mIoU from 0.6310 to 0.9129, over the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。