arXiv:2608.05505cs.CV2026-08

提出动态像素基准,测试模型能否精准预测未来真实画面。

DynaPix: Can Vision-Language Models Identify the Exact Future?

论文配图:DynaPix: Can Vision-Language Models Identify the Exact Future?
图 1 · 摘自论文原文
  • 构建可验证的未来预测评测集,基于物理模拟器生成精确答案
  • 模型在事件标记时表现良好,但对时间标记的预测接近随机
  • 适合研究视觉语言模型时空推理能力的学者使用

在真实场景中行动需要预知其确切的未来状态,而非仅一个合理画面。当前评估常接受文字描述或逼真图像,导致预测结果无法与真实状态对比。我们提出DynaPix(动态像素)基准,使未来预测可被验证:给定一段视频片段(事件前停止)和关于后续时刻的问题,模型需从相近候选图中选出真实未来图像,或从大图库中检索。场景源自物理模拟器,正确图像及其时间点完全可知,错误选项刻意设计得高度相似。当可见事件标记目标时刻时,模型表现尚可;但仅以时间流逝标记时,性能接近随机。图库搜索更难,真实图像极少排第一。人类能较好处理时间标记任务,说明问题出在模型而非问题设计。用模拟器真实记录训练,可部分修复该缺陷,但对长时序任务改善有限。因此,DynaPix揭示了模型在时间锚定上的显著差距:对事件的关联远强于对时间本身的理解。

原文摘要 · Abstract (English)

Acting in a physical scene requires knowing its real later state, not a plausible one. Current evaluations often accept words or a realistic-looking image, so the predicted state is never checked against the true one. We introduce DynaPix (Dynamic Pixels), a benchmark that makes prediction checkable. Given a video clip that stops before a key event and a question about a later moment, a model must pick the true future image from close candidates or a large gallery. The scenes come from a physics simulator, so the correct image and its time are known exactly and the wrong options are deliberately similar. Models often succeed when a visible event marks the target moment, but are near chance when only elapsed time marks it. Gallery search is harder still, as the true image rarely ranks first. People handle the elapsed-time items well, so the difficulty lies with the models, not the questions. Training on scene accounts drawn from the simulator's true record, not a teacher's guess, repairs much of this but not the longer elapsed time case. DynaPix thus exposes a temporal-anchoring gap: models attach a prediction to an event far better than to time itself.

视觉推理时间预测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。