arXiv:2607.19228cs.CV2026-07

实时理解动态场景中物体的4D变化,实现几何与实例的一致性重建。

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

论文配图:IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer
图 1 · 摘自论文原文
  • 通过因果时空建模,逐步融合视频流中的几何与对象身份信息。
  • 在14.7万帧数据集上实现长序列动态场景的高精度重建与跟踪。
  • 适合需要长期动态环境理解的自动驾驶与机器人应用。

现实世界的空间智能要求智能体从连续视频流中理解场景,其中物体随时间移动、持续存在、消失或重新出现。尽管近期空间基础模型实现了可泛化的前馈3D重建,但大多数流式方法仍以几何为中心,缺乏时序一致的对象级理解。现有语义重建和3D感知视觉-语言方法大多依赖外部提取的2D语义线索或松散耦合的几何输入,限制了长动态场景中统一的几何-实例学习。本文提出IGGT4D,一种用于在线4D场景理解的流式实例锚定几何变换器。IGGT4D按顺序处理视频帧,通过因果时空建模复用历史上下文,并增量更新相机运动、几何与对象身份的统一表示。这使得在动态环境中实现长序列前馈重建并保持几何-实例一致性成为可能。为解决高质量4D监督数据缺失问题,我们进一步构建了InsScene4D-147K,一个大规模数据集,涵盖真实/合成及静态/动态场景,包含RGB图像、深度图、位姿和由自动化几何引导标注流程生成的时间一致实例掩码。在3D重建、位姿估计、实例空间追踪和开放词汇分割任务上的实验表明,IGGT4D优于现有流式基线,同时保持对长动态序列的可扩展在线推理能力。

原文摘要 · Abstract (English)

Real-world spatial intelligence requires agents to understand scenes from continuous video streams, where objects move, persist, disappear, and reappear over time. While recent spatial foundation models have enabled generalizable feed-forward 3D reconstruction, most streaming methods remain geometry-centric and lack temporally consistent object-level understanding. Meanwhile, existing semantic reconstruction and 3D-aware vision-language methods largely rely on externally extracted 2D semantic cues or loosely coupled geometry inputs, limiting unified geometry-instance learning in long dynamic scenes. In this paper, we propose IGGT4D, a streaming instance-grounded geometry Transformer for online 4D scene understanding. IGGT4D processes video frames sequentially, reuses historical context through causal spatial-temporal modeling, and incrementally updates a unified representation of camera motion, geometry, and object identity. This enables long-sequence feed-forward reconstruction with geometry-instance consistency in dynamic environments. To address the lack of high-quality 4D supervision, we further construct InsScene4D-147K, a large-scale dataset spanning real/synthetic and static/dynamic scenes, with RGB images, depth, poses, and temporally consistent instance masks generated by an automated geometry-guided annotation pipeline. Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.

4D重建实例感知流式处理时空建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。