arXiv:2607.22534cs.CVcs.AI2026-07

让3D动态重建学会物体整体运动规律,更真实地还原动态场景。

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

论文配图:SM4RT: Learning Structured Motion Geometry for 4D Reconstruction
图 1 · 摘自论文原文
  • 将运动建模为刚体变换的基向量,实现结构化动态感知
  • 在单目视频上端到端重建3D几何与运动,精度优于现有方法
  • 适合需要精确动态建模的机器人、虚拟现实应用

几何基础模型(GFMs)显著提升了单目3D重建能力,但将其扩展至4D动态理解仍面临根本性挑战。现有运动感知方法(如稀疏跟踪、密集点对流场)通常将运动视为独立点位移,忽略了物理运动的结构性。实际上,真实物体通常遵循刚体运动规律,点之间呈现集体运动而非孤立移动。运动本身具有几何结构:物体经历由SE(3)控制的一组刚体变换,而非无结构的点对位移。基于此,我们提出SM4RT——一种面向端到端3D重建与结构化运动感知的结构化运动4D重建变压器。SM4RT引入运动-结构表示,将场景运动分解为一组紧凑的运动基,每个基以SE(3)中的6维旋转变换时间序列表示。通过稀疏、时共享的像素级分配权重恢复稠密场景运动,确保同一物体上的点共享相同的刚体运动轨迹。SM4RT采用并行运动几何编码器与解码器,在单次前向传播中联合推断3D几何、世界坐标系运动与场景运动结构,实现高保真的运动重建。

原文摘要 · Abstract (English)

Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge. Most existing motion perception methods (e.g., sparse tracking, dense point-wise flow) treat motion as independent point-wise displacements, ignoring the structured nature of physical motion. However, real-world objects usually obey rigid-body kinematics, and points thus usually move collectively, not in isolation. Motion itself possesses geometric structure: physical objects undergo a set of rigid-body transformations governed by SE(3), rather than unstructured point-wise displacements. Building on this insight, we propose SM4RT, a Structured Motion 4D Reconstruction Transformer for end-to-end 3D reconstruction and structured motion perception. SM4RT introduces Structure-of-Motion to represent scene dynamics, where scene motion is decomposed into a compact set of motion bases, each represented as a temporal sequence of 6D twists in SE(3). Dense scene motion is then recovered by sparse, time-shared per-pixel assignment weights over these bases, ensuring points on the same object share a common rigid-body motion trajectory. SM4RT introduces a parallel motion geometry encoder and decoder that jointly infer 3D geometry, world-coordinate motion, and scene kinematic structure in a single forward pass from monocular RGB video. SM4RT achieves strong motion reconstruction performance while preserving the geometric structure of scene motion.

4D重建运动建模刚体运动视觉几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。