让大模型通过视觉化假想场景来增强推理能力
Einstein World Models

- 在推理过程中引入视觉时序推演,生成可检视的假想场景
- 使大模型能进行文本难以支持的复杂逻辑推演
- 适合需要深度想象和假设验证的研究者与开发者
智能是否需要理解直接经验之外的现象?我们怀疑某些复杂思维无法仅通过语言捕捉。本文关注的是:可视化反事实事件能否作为语言的补充,促进复杂思考。我们探究大语言模型(LLMs)是否可通过训练利用此类视觉化机制以提升推理能力。为此,提出爱因斯坦世界模型(Einstein World Models, EWM)。EWM 是一种基于大模型的推理系统蓝图,将视觉-时间推演嵌入推理轨迹中,使模型能够以文本难以支持的方式进行推理。在 EWM 中,大模型调用一个世界模块(非世界模型),生成所考虑场景的短时序推演。返回的推演结果不作为答案,而是作为可检视的假设,支持后续推理。EWM 将大模型的工具调用能力(如网络搜索或代码执行)拓展至视觉思想实验领域。
原文摘要 · Abstract (English)
Does intelligence require the ability to reason about phenomena beyond direct experience? It is natural to suspect that some complex thought cannot be captured through language alone. However, of particular concern to this work, is whether visualising counterfactual events can complement language as a mechanism for complex thought. We ask whether LLMs can be trained to utilise such visualisation mechanisms, in a way that benefits their reasoning abilities. Motivated by this question, we propose Einstein World Models. EWMs are a blueprint for LLM-based reasoning systems that place visual-temporal rollouts inside the reasoning trace, allowing them to reason in ways that text alone may not support well. In an EWM, the LLM calls a world-module (not to be confused with a world model), to produce short rollouts of scenes under consideration. The returned rollout is treated not as the answer, but as an inspectable hypothesis that can support later reasoning. Einstein World Models extend the capability of LLMs for tool calling (such as web search or code execution), into the domain of visual thought experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。