提出4DVLT框架,让多视角视频中物体的动态轨迹与语言指令精准对齐。
4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking

- 用4D状态图建模物体运动,通过语义引导实现跨视角连续追踪
- 在129.4万问答数据上达到62.68的Top1精度,超越基线19.62点
- 适合需要精确理解动态场景中物体身份与运动的智能系统开发者
4D动态场景理解需将语言指令与持续的物体世界线(worldline)对齐,以关联身份、度量3D运动及同步多视角2D投影。现有方法仅捕捉该结构的部分特征:大型多模态模型虽能利用丰富视觉证据推理,但极少保持度量拓扑;而视觉语言追踪仍局限于碎片化2D或3D输出及局部延续。为此,我们提出4DVLT——一种以世界线为中心的指令条件4D动态场景理解任务,并构建包含129.4万问题-答案对、64.7万目标实体、851个场景及9类推理型查询的基准数据集Instruct-4D。为应对此任务,我们设计4DTrack,将指令条件追踪建模为图条件下的世界线推断,结合对象中心4D状态图、度量引导路由、双向解码与运动学校准。在Instruct-4D上,4DTrack-Qwen3.5-9B达到62.68的$ ext{TGA}_{ ext{Top1}}$,优于最佳适配的视觉语言追踪基线19.62点。结果表明,以世界线为中心的建模显著提升目标定位与恢复的世界线质量。
原文摘要 · Abstract (English)
4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view 2D projections. Existing paradigms capture only part of this structure: large multimodal models reason over rich visual evidence but rarely preserve metric topology, while vision-language tracking remains tied to fragmented 2D or 3D outputs and local continuation. We therefore introduce \textbf{4DVLT}, a worldline-centered task for instruction-conditioned 4D dynamic scene understanding in fully observed multi-view video, and \textbf{Instruct-4D}, a benchmark with 129.4K question-answer pairs, 64.7K target entities, 851 scenes, and 9 reasoning-oriented query types. To address this setting, we present \textbf{4DTrack}, which casts instruction-conditioned tracking as graph-conditioned worldline inference through an object-centric 4D state graph, metric-guided routing, bidirectional decoding, and kinematic calibration. On Instruct-4D, 4DTrack-Qwen3.5-9B reaches 62.68 $\mathrm{TGA}_{\mathrm{Top1}}$ and surpasses the best adapted VLT baseline by 19.62 points. These results show that worldline-centered modeling improves both target grounding and recovered worldline quality. The project page is available at https://github.com/mikubaka88/4DVLT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。