用3D高斯点云实现多人多物体动态场景的高质量渲染
Rendering Multi-Human and Multi-Object with 3D Gaussian Splatting
- 分步构建:先对每个物体独立融合多视角信息,再全局建模交互关系
- 在复杂遮挡下仍能保持视觉一致性,生成真实接触效果
- 适合机器人仿真、VR/AR等需要高保真数字孪生的场景
从稀疏视角输入重建多个相互作用的人体与物体的动态场景是一项关键但极具挑战的任务,对创建用于机器人和虚拟现实/增强现实的高保真数字孪生至关重要。我们称此问题为多人多物体(MHMO)渲染,其面临两大难题:在严重相互遮挡下为各实例建立视图一致的表示,以及显式建模由交互产生的复杂组合依赖关系。为此,我们提出MM-GS,一种基于3D高斯点云的新型分层框架。方法首先通过每实例多视角融合模块,聚合所有可用视角的信息,建立鲁棒且一致的实例表示;随后,场景级实例交互模块在全局场景图上推理所有参与者之间的关系,优化其属性以捕捉细微的交互效应。在挑战性数据集上的大量实验表明,该方法显著优于强基线,生成具有高保真细节和合理实例间接触状态的最新成果。
原文摘要 · Abstract (English)
Reconstructing dynamic scenes with multiple interacting humans and objects from sparse-view inputs is a critical yet challenging task, essential for creating high-fidelity digital twins for robotics and VR/AR. This problem, which we term Multi-Human Multi-Object (MHMO) rendering, presents two significant obstacles: achieving view-consistent representations for individual instances under severe mutual occlusion, and explicitly modeling the complex and combinatorial dependencies that arise from their interactions. To overcome these challenges, we propose MM-GS, a novel hierarchical framework built upon 3D Gaussian Splatting. Our method first employs a Per-Instance Multi-View Fusion module to establish a robust and consistent representation for each instance by aggregating visual information across all available views. Subsequently, a Scene-Level Instance Interaction module operates on a global scene graph to reason about relationships between all participants, refining their attributes to capture subtle interaction effects. Extensive experiments on challenging datasets demonstrate that our method significantly outperforms strong baselines, producing state-of-the-art results with high-fidelity details and plausible inter-instance contacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。