让AI看懂视频中物体如何随时间变化,提升复杂推理能力。
ChronoVision: Temporal Reasoning via Latent State Reconstruction

- 通过重建最终状态的潜在图像,对齐视觉逻辑与模型理解。
- 在新数据集Vbvr-VQA上达到74.8%(域内)和71.6%(域外)准确率。
- 适合需要多步时序推理的视频理解任务,如因果推断与轨迹预测。
多模态大语言模型在被动感知上表现优异,但在需要多步时间推理的复杂视觉认知任务中表现下降,这主要源于语言推理的固有模糊性,难以准确描述连续视觉变化。为此,我们提出ChronoVision,一种将视觉逻辑与潜在图像对齐的多模态框架。在监督微调阶段,重构式视觉头预测最终变换状态的潜在表示,而ROI注意力定位模块通过语义跨度查询聚焦关键视觉证据。后训练阶段,采用强化学习结合隐式过程对齐机制,由复合奖励函数评估结果正确性、潜在过程一致性及无监督视觉关注。此外,我们构建了新数据集Vbvr-VQA,将视频推理转化为严格的图像排序任务以评估时间追踪能力。实验表明,ChronoVision在Vbvr-VQA上实现74.8%(域内)和71.6%(域外)准确率,并在极具挑战性的跨域基准IntPhys2上取得55.0%准确率,性能达到当前最佳。
原文摘要 · Abstract (English)
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, which often fails to accurately articulate continuous visual transformations. To address this, we propose ChronoVision, a multimodal framework designed to align visual logic with latent imagery. During supervised fine-tuning, a Reconstructive Visual Head predicts the latent representation of the final transformed state, while an ROI Attention Locating module focuses the model on key visual evidence via semantic span queries. In post-training, we apply reinforcement learning with an implicit process grounding mechanism, guided by a composite reward function that evaluates outcome correctness, latent process alignment, and unsupervised visual focus. Furthermore, we introduce Vbvr-VQA, a novel dataset that evaluates temporal tracking by reformulating video reasoning into a strict image-ordering task. Experiments demonstrate that ChronoVision achieves state-of-the-art performance on Vbvr-VQA with 74.8% in-domain and 71.6% out-of-domain accuracy, alongside a strong 55.0% accuracy on IntPhys2, a highly challenging cross-domain benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。