arXiv:2512.00547cs.CV2025-12

针对多人多物动态场景,提出融合生成模型与高斯溅射的重建方法。

Asset-Driven Sematic Reconstruction of Dynamic Scene with Multi-Human-Object Interactions

  • 结合生成模型与高斯溅射,分步优化物体与人体结构。
  • 在严重遮挡下仍保持几何一致性,提升多视角与时间连续性。
  • 适合需要高保真动态场景重建的AR/VR与智能体应用。

真实世界中的人工环境高度动态,涉及多人与周围物体的复杂交互。尽管3D几何建模对AR/VR、游戏和具身智能等应用至关重要,但因运动模式多样和频繁遮挡,该领域仍研究不足。3D高斯溅射(GS)在快速优化结构并生成高质量表面几何方面表现优异,但极少有基于GS的方法处理多人多物场景,主要受限于上述挑战。单目设置下,仅依赖GS渲染损失优化时,严重遮挡下的结构一致性更难维持。为此,本文提出一种混合方法:1)利用3D生成模型生成场景元素的高保真网格;2)通过语义感知变形(刚体对象刚性变换,人体基于线性骨骼绑定的形变)将变形后的高保真网格映射至动态场景;3)基于GS优化各元素以进一步精修其空间对齐。该方法在严重遮挡下仍能保持物体结构稳定,实现多视角与时间一致的几何重建。我们选用HOI-M3数据集进行评估,因其是目前唯一包含多人多物动态交互的数据集。实验表明,本方法在表面重建质量上优于现有最先进方法。

原文摘要 · Abstract (English)

Real-world human-built environments are highly dynamic, involving multiple humans and their complex interactions with surrounding objects. While 3D geometry modeling of such scenes is crucial for applications like AR/VR, gaming, and embodied AI, it remains underexplored due to challenges like diverse motion patterns and frequent occlusions. Beyond novel view rendering, 3D Gaussian Splatting (GS) has demonstrated remarkable progress in producing detailed, high-quality surface geometry with fast optimization of the underlying structure. However, very few GS-based methods address multihuman, multiobject scenarios, primarily due to the above-mentioned inherent challenges. In a monocular setup, these challenges are further amplified, as maintaining structural consistency under severe occlusion becomes difficult when the scene is optimized solely based on GS-based rendering loss. To tackle the challenges of such a multihuman, multiobject dynamic scene, we propose a hybrid approach that effectively combines the advantages of 1) 3D generative models for generating high-fidelity meshes of the scene elements, 2) Semantic-aware deformation, \ie rigid transformation of the rigid objects and LBS-based deformation of the humans, and mapping of the deformed high-fidelity meshes in the dynamic scene, and 3) GS-based optimization of the individual elements for further refining their alignments in the scene. Such a hybrid approach helps maintain the object structures even under severe occlusion and can produce multiview and temporally consistent geometry. We choose HOI-M3 for evaluation, as, to the best of our knowledge, this is the only dataset featuring multihuman, multiobject interactions in a dynamic scene. Our method outperforms the state-of-the-art method in producing better surface reconstruction of such scenes.

动态场景多人体交互高斯溅射3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。