从单目视频重建可操控的3D动态场景,支持机器人任务执行
ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting

- 分离机器人、物体与背景的3D高斯子场,构建场景图结构
- 通过动作-技能交替逻辑生成伪真值轨迹,提升重建稳定性
- 重建结果物理一致且适配仿真,可用于机器人策略学习
从真实世界观测中重建动态交互式3D场景仍是计算机视觉与机器人领域的基础挑战。尽管3D高斯点阵技术已实现高保真静态重建,但面对带关节机器人和可操作物体的交互环境时,复杂的接触交互与突发姿态变化仍带来困难。为此,我们提出ManiSplat,一种统一框架,直接从单目自视角机器人视频中重建可控制、解耦的高斯数字孪生。方法引入图结构解耦表示,将机器人、物体与背景分离为独立优化的高斯子场,并组织于场景图中。为保证稳定性,设计任务导向时空对齐模块,利用操作任务固有的逻辑——动作与技能阶段交替——构造准确伪真值轨迹。最后,联合光度-几何优化确保重建场景在时间上连贯、物理一致且可直接用于仿真。大量实验证明,该方法能以高保真度重建驱动交互的动态场景,有效支持下游机器人任务与策略学习。
原文摘要 · Abstract (English)
Reconstructing dynamic and interactive 3D scenes from real-world observations remains a fundamental challenge in computer vision and robotics. While recent advances in 3D Gaussian Splatting have enabled high-fidelity static reconstruction, extending it to interactive environments with articulated robots and manipulable objects remains difficult due to complex contact interactions and abrupt pose changes. To address these challenges, we introduce ManiSplat, a unified framework that reconstructs controllable and decoupled Gaussian digital twins directly from monocular ego-view robotic videos. Our method introduces a Graph-Structured Disentangled Representation that separates the robot, objects, and background into independently optimizable Gaussian subfields organized within a scene graph. To ensure stability, we propose a Task-Oriented Spatio-Temporal Alignment module that leverages the inherent logic of manipulation tasks-alternating between Motion and Skill phases-to construct accurate pseudo-ground-truth trajectories. Finally, a joint photometric-geometric optimization ensures the reconstructed scenes are temporally coherent, physically consistent, and simulation-ready. Extensive experiments demonstrate that our approach reconstructs interaction-driven dynamic scenes with high fidelity and controllability, effectively supporting downstream robotic tasks and policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。