arXiv:2412.04457cs.CV2024-12被引 7

对比多种单目动态场景高斯点云方法,发现真实数据复杂度掩盖了性能差异。

Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules

  • 系统分类并对比多类单目动态高斯点云方法的运动建模方式
  • 合成数据中方法排序清晰,但真实场景中性能差异被复杂度掩盖
  • 所有方法均快速渲染,但优化过程易崩溃,存在本质脆弱性

高斯点云方法正成为将多视角图像转换为可进行视图合成的场景表示的热门选择。尤其值得关注的是,仅用单目输入实现动态场景的视图合成,这是一个病态且极具挑战的问题。该领域进展迅速,已有多篇论文声称效果最佳,但不可能全部成立。本文系统组织、基准测试并分析了多种基于高斯点云的方法,提供了此前研究缺乏的公平比较。我们使用多个现有数据集和一个新设计的合成数据集,该数据集旨在分离影响重建质量的关键因素。我们系统地将高斯点云方法归类为特定的运动表示类型,并量化其差异对性能的影响。实证结果表明,在合成数据中方法的排名是明确的,但在真实世界数据中,复杂度已压倒性能差异。此外,所有基于高斯的方法虽具有快速渲染的优势,但优化过程存在显著脆弱性。我们总结实验发现,为该活跃问题的研究提供指导。

原文摘要 · Abstract (English)

Gaussian splatting methods are emerging as a popular approach for converting multi-view image data into scene representations that allow view synthesis. In particular, there is interest in enabling view synthesis for dynamic scenes using only monocular input data -- an ill-posed and challenging problem. The fast pace of work in this area has produced multiple simultaneous papers that claim to work best, which cannot all be true. In this work, we organize, benchmark, and analyze many Gaussian-splatting-based methods, providing apples-to-apples comparisons that prior works have lacked. We use multiple existing datasets and a new instructive synthetic dataset designed to isolate factors that affect reconstruction quality. We systematically categorize Gaussian splatting methods into specific motion representation types and quantify how their differences impact performance. Empirically, we find that their rank order is well-defined in synthetic data, but the complexity of real-world data currently overwhelms the differences. Furthermore, the fast rendering speed of all Gaussian-based methods comes at the cost of brittleness in optimization. We summarize our experiments into a list of findings that can help to further progress in this lively problem setting.

动态场景高斯点云单目重建方法对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。