让AI看视频时能‘回头再看’,提升空间推理准确率
Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

- 先基于原始视频形成空间假设,再通过合成新视角视频验证或修正
- 在VSI-Bench和STI-Bench上使开源模型性能接近闭源顶尖水平
- 无需训练、不改架构,只需生成互补视角视频辅助推理
从第一人称视频进行空间推理本质上具有挑战性,因可观测证据受限于相机轨迹。现有方法依赖单次推理,迫使模型依靠语义先验解决几何模糊问题,而非可验证证据。我们提出「先推理,再重推理」(ReRe)框架,一种无需训练的推理时机制:在推理阶段,多模态大模型基于原始视频形成空间假设;在重推理阶段,通过观察由预测3D几何生成的合成新视角视频来验证或修正假设。为此,我们设计了几何转视频管道,生成具有更高、斜向视角且覆盖整个场景的互补新视图,同时保持大模型原有视频输入接口不变。在VSI-Bench和STI-Bench上的大量实验表明,ReRe显著提升开源多模态大模型性能,使其媲美闭源最先进水平。
原文摘要 · Abstract (English)
Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory. Existing methods rely on single-turn inference, forcing models to resolve geometric ambiguity through semantic priors rather than verifiable evidence. We argue that spatial reasoning should be revisitable: conclusions formed under limited evidence should remain open to revision when complementary viewpoints become available. Building on this insight, we propose Reason, then Re-reason (ReRe), a training-free, inference-time framework with two phases: in the Reason Phase, an MLLM forms a spatial hypothesis from the original video; in the Re-reason Phase, it verifies or revises the hypothesis by observing a synthesized novel-view video. To enable effective cross-view revisiting, we design a Geometry-to-Video pipeline that renders strategically complementary novel views from predicted 3D geometry. These views feature an elevated, oblique perspective with scene-spanning coverage, while preserving the MLLM's native video interface without architectural modifications. Extensive evaluations on VSI-Bench and STI-Bench demonstrate that ReRe substantially boosts open-source MLLMs to rival proprietary state-of-the-art performance. Project page: https://zhenjiemao.github.io/ReRe/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。