arXiv:2606.28820cs.CV2026-06

从单目视频重建人与物体动态交互场景,分离人体、物体和背景的运动模型。

CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video

论文配图:CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video
图 1 · 摘自论文原文
  • 将人体、物体和背景分三路建模,分别用姿态先验、轨迹驱动和平面约束
  • 在HOSNeRF和NeuMan数据集上提升人-物交互重建精度与真实感
  • 适合需要高保真人-物交互渲染的虚拟现实与动画制作

从单目视频重建人与物体交互的动态场景极具挑战性,因人体、被操作物体与背景遵循不同运动模型却共享同一像素。现有动态辐射场与高斯溅射方法常混淆各组件,导致物体运动泄露至人体或静态场景,且单目人体重建在罕见观测区域仍受约束不足。本文提出CoGS,一种用于单目人-物场景重建的组合式高斯溅射框架。CoGS将视频分解为三个协同分支:基于完整姿态先验初始化的人体,由估计物体轨迹驱动的刚性物体场,以及在可用时以弱场景平面原型正则化的静态场景场。六阶段优化流程先独立稳定人体与物体,再通过全图监督、可见性感知人体锚定、物体轮廓与运动约束及延迟场景正则化融合三者。该设计使各组件专注自身几何与运动,同时允许光度证据修正最终复合结果。在HOSNeRF与NeuMan上的实验表明,CoGS在人-物交互重建与野外人-场景渲染方面均实现更强保真度与感知质量,涵盖全帧与聚焦人体的评估。代码将在发表后公开。

原文摘要 · Abstract (English)

Reconstructing dynamic human--object interaction scenes from monocular video is difficult because the human, manipulated object, and background obey different motion models while sharing the same pixels. Existing dynamic radiance-field and Gaussian-splatting methods often entangle these components, causing object motion to leak into the human or static scene, and monocular human reconstruction remains underconstrained in regions that are rarely observed. We present CoGS, a compositional Gaussian-splatting framework for monocular human--object scene reconstruction. CoGS decomposes the video into three coordinated branches: an articulated human initialized from a complete canonical prior, a rigid object field driven by an estimated object trajectory, and a static scene field regularized by weak scene-only planar primitives when available. A six-stage optimization schedule first stabilizes the human and object independently, then fuses them with the scene under full-image supervision, visibility-aware human anchoring, object silhouette and motion constraints, and delayed scene regularization. This design keeps each component responsible for its own geometry and motion while allowing photometric evidence to correct the final composite. Experiments on HOSNeRF and NeuMan show that CoGS improves both human--object interaction reconstruction and in-the-wild human--scene rendering, achieving stronger fidelity and perceptual quality across full-frame and human-focused evaluations. Code will be released upon publication.

三维重建动态场景高斯溅射人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。