构建可控合成数据集,评估3D/4D高斯点云的重建与压缩性能。
M$^3$ISR: A Multi-Modal Multi-View Benchmark for 3D/4D Gaussian Splatting and Feedforward Compression

- 设计共享中心相机结构,隔离视角变化影响,实现可控评估。
- 包含25个场景、6路1080p同步视图及多模态标注,支持多任务测试。
- 首次提出前馈压缩任务,提供速率-失真基准,适合算法优化研究者。
高保真自由视角视频与交互渲染日益依赖显式高斯表示,但实际部署受限于表示规模、动态更新与计算成本。现有多视角视频基准虽提供真实采集内容,却难以分离相机几何、表示效率与时间冗余的影响。本文提出M$^3$ISR,一个面向3D/4D高斯点云(3DGS/4DGS)的受控合成基准。包含25个场景(五类室内外场景),两种相机/运动配置,六路同步1080p视图,以及密集真值标注:RGB、相机参数、深度、语义与实例分割、静态-动态掩码。共享中心相机设计旨在分离视角变化,支持新视角合成与表示效率的可控评估。基准分为五个互补赛道:3DGS重建、4DGS重建、4DGS流传输、3DGS压缩、4DGS压缩。基线结果表明,静态重建质量差异小,但表示存储量差异显著;流传输方法的训练或重建成本远高于离线动态重建基线。此外,定义了3DGS与4DGS的前馈压缩任务,提供参考速率-失真公式与初步基线评估。该基准旨在作为系统性研究高斯基自由视角视频重建、压缩与传输的受控互补平台。
原文摘要 · Abstract (English)
High-fidelity free-viewpoint video (FVV) and interactive rendering increasingly rely on explicit Gaussian representations, yet practical deployment remains constrained by representation size, dynamic updates, and computational cost. Existing multi-view video benchmarks provide valuable real-captured content, but they make it difficult to isolate the effects of controlled camera geometry, representation efficiency, and temporal redundancy. We introduce M$^3$ISR, a controlled synthetic benchmark for 3D and 4D Gaussian Splatting (3DGS/4DGS). The benchmark contains 25 scenes from five indoor and outdoor scene groups, two camera/motion configurations, six synchronized 1080p views, and dense ground-truth annotations including RGB, camera parameters, depth, semantic and instance segmentation, and static--dynamic masks. The shared-center camera design intentionally isolates angular view variation and enables controlled evaluation of novel-view synthesis and representation efficiency. We organize M$^3$ISR into five complementary tracks covering 3DGS synthesis, 4DGS synthesis, 4DGS streaming, 3DGS compression, and 4DGS compression. Representative baseline results show small differences in static reconstruction quality but substantial differences in representation storage, while the evaluated streaming methods exhibit substantially higher reported training or reconstruction cost than the corresponding offline dynamic reconstruction baselines. We further define feedforward compression tasks for 3DGS and 4DGS and provide reference rate--distortion formulations and preliminary baseline evaluations. The benchmark is intended as a controlled and complementary testbed for systematic study of Gaussian-based FVV reconstruction, compression, and streaming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。