构建首个预测智能评测数据集,推动大模型理解未来事件。
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
- 设计新数据集FSU-QA,专用于评估视觉语言模型的预见能力。
- 小模型微调后超越大模型,证明数据集能显著提升预见推理。
- 可用于评估世界模型生成内容的语义一致性,适合未来场景研究者。
本文提出将预见智能定义为预判和解读未来事件的能力,该能力对自动驾驶等应用至关重要,但当前研究普遍忽视。为此,我们构建了专为激发和评估预见智能而设计的新视觉问答数据集FSU-QA。基于此,首次系统评估了主流视觉语言模型在预见任务中的表现,发现现有模型仍难以有效推理未来情境。除作为基准外,FSU-QA还可通过衡量世界模型生成预测的语义连贯性来评估其质量,具体以增强后模型性能提升为量化指标。实验表明,即使微调小型模型,其在预见推理上也远超大型先进模型。此外,我们使用视觉语言模型作为代理裁判,检验世界模型生成结果的语义一致性,并通过随机对照实验验证评估方法的有效性。这些成果确立了FSU-QA作为下一代可预见模型研发的坚实基础。
原文摘要 · Abstract (English)
In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applications such as autonomous driving, yet largely overlooked by existing research. To bridge this gap, we introduce FSU-QA, a new Visual Question-Answering (VQA) dataset specifically designed to elicit and evaluate Foresight Intelligence. Using FSU-QA, we conduct the first comprehensive study of state-of-the-art Vision-Language Models (VLMs) under foresight-oriented tasks, revealing that current models still struggle to reason about future situations. Beyond serving as a benchmark, FSU-QA also enables the assessment of world models by measuring the semantic coherence of their generated predictions, quantified through performance gains when VLMs are augmented with such outputs. Our experiments further demonstrate that FSU-QA can effectively enhance foresight reasoning: even small VLMs fine-tuned on FSU-QA surpass much larger, advanced models by a substantial margin. Together, these findings position FSU-QA as a principled foundation for developing next-generation models capable of truly anticipating and understanding future events. Furthermore, beyond model performance, we examine whether WM-generated predictions remain semantically consistent by using VLM-based proxy judges, and validate this evaluation protocol through shuffled control experiments. Fine-tuning models on FSU-QA leads to substantial improvements in foresight understanding, demonstrating the dataset's effectiveness and offering a principled foundation for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。