让AI从猜词变成懂物理因果,提升真实世界决策能力
Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners

- 用多阶段验证构建高保真因果评估基准
- 百万级推理轨迹数据使模型物理预测准确率提升36.3%
- 适合追求真实物理推理能力的智能体研究者
当前具身视觉-语言规划评测倾向于奖励语言统计模式而非物理因果推理,导致模型仅模仿文本规律而忽略物理逻辑。本文提出Causal-Plan-Bench,一个通过多阶段验证构建的高保真诊断套件,用于评估四种因果维度下的具身规划能力。同时构建了包含一百万条显式推理轨迹的Causal-Plan-1M数据集,基于四阶段标注流程生成于第一视角视频。实验表明,主流模型仍难以展现真正物理自主性,Gemini 3 Pro在该基准上仅达38.18分。相比之下,基于Qwen3-VL-8B的Causal Planner通过特定训练方案,内化物理逻辑,在域内表现优异且跨基准泛化能力强。其性能遵循因果缩放定律:将因果训练数据扩展至百万规模,相对增益达36.3%(从33.22提升至45.28)。本工作为推动智能体从表面词序预测转向物理因果推理提供了具体路径。
原文摘要 · Abstract (English)
Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning. This rewards models that mimic statistical language priors rather than track causal dependencies, reducing physical planning to shallow sequence modeling. We argue that reliable physical autonomy requires a shift from linguistically grounded token prediction toward physically grounded causal reasoning. To this end, we introduce Causal-Plan-Bench, a high-fidelity diagnostic suite curated through multi-stage verification to evaluate embodied planning across four causal dimensions. We also construct Causal-Plan-1M, a million-scale corpus of explicit reasoning traces produced by a four-stage annotation pipeline over egocentric videos. Extensive evaluation shows that leading models still struggle to demonstrate genuine physical agency, with Gemini 3 Pro reaching only 38.18 on our benchmark. In contrast, our training recipe enables Causal Planner, built on Qwen3-VL-8B, to internalize physical logic for more accurate next-state estimation. The model achieves strong in-domain performance and cross-benchmark generalization, and reveals a Causal Scaling Law: scaling causal training data to one million instances yields a 36.3% relative gain, from 33.22 to 45.28. Overall, our work provides a concrete step toward turning agents from superficial token predictors into physically grounded causal reasoners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。