让大模型持续追踪动态物体,提升对时空变化的理解能力。
DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs

- 用几何引导的轨迹可视化和动态痕迹图,连续追踪物体运动
- 在多个基准上超越现有方法,实现4D推理最佳性能
- 适合需要精确动态场景理解的任务,如机器人交互
4D时空推理(联合建模三维空间结构与时间演化)对于理解动态世界和实现具身交互至关重要。当前多模态大模型虽在静态场景理解和粗粒度4D任务中表现良好,但在连续动态场景感知方面仍存在明显不足,尤其难以持续追踪动态物体证据以实现连贯的4D时空推理。其主要原因是依赖稀疏帧级观测,割裂了连续动态线索,导致模型无法区分真实物体运动与相机引起的视差运动。受人类在视角变化下持续追踪动态线索的启发,我们提出无需训练的DynTrace框架,包含两个互补组件:动态轨迹可视化(DTV)将世界坐标轨迹投影至图像平面,提供几何引导的视觉先验,以分离真实运动与相机伪运动;动态痕迹令牌(DT-Token)按动态痕迹图(DTG)组织,追踪物体级动态线索、演变过程及关键时刻,保持连续的动态物体证据。二者共同使大模型具备基于几何先验与结构化时空轨迹的连续动态证据追踪能力。DynTrace在开源大模型上持续提升性能,在Dyn-Bench、VLM4D和DSI-Bench三个基准上达到最先进水平,验证了持续追踪动态物体证据对鲁棒4D时空推理的重要性。
原文摘要 · Abstract (English)
4D spatio-temporal reasoning, jointly modeling 3D spatial structure and temporal evolution, is essential for understanding dynamic worlds and enabling embodied interaction. While current Multimodal Large Language Models (MLLMs) show strong capabilities in static scene understanding and coarse-grained 4D tasks, they still have notable limitations in continuous dynamic scene perception, especially in tracking dynamic object evidence for coherent 4D spatio-temporal reasoning. This shortcoming stems mainly from relying on sparse frame-level observations, fragmenting continuous dynamic cues and leaving models unable to disentangle genuine object dynamics from camera-induced apparent motion. Inspired by humans tracking dynamic cues while compensating for viewpoint changes, we propose DynTrace, a training-free framework for 4D spatio-temporal reasoning with two complementary components. Dynamic Trajectory Visualization (DTV) reprojects world-coordinate trajectories onto the image plane, providing geometry-informed visual priors that disentangle genuine object dynamics from camera-induced apparent motion. Meanwhile, the Dynamic Trace Token (DT-Token), organized into a Dynamic Trace Graph (DTG), tracks object-level dynamic cues, trace evolution, and key moments, maintaining continuous dynamic object evidence for coherent 4D reasoning. Together, these two components equip MLLMs with continuously tracked dynamic object evidence, grounded in geometry-informed visual priors and structured spatio-temporal traces. DynTrace consistently improves open-source MLLMs, achieving state-of-the-art results on Dyn-Bench, VLM4D, and DSI-Bench, validating the importance of tracking dynamic object evidence for robust 4D spatio-temporal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。