arXiv:2509.06951cs.ROcs.CV2025-09被引 70

F1模型通过预测未来视觉状态,让AI更智能地规划动作。

F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions

  • 用视觉前瞻生成来指导决策,把动作选择变成预见未来的逆动力学问题
  • 在超过33万条轨迹上训练,136个任务中成功率显著提升
  • 适合需要长期规划和动态适应的机器人或AI代理场景

在动态视觉环境中执行语言指令的任务仍是具身AI的核心挑战。现有视觉-语言-动作(VLA)模型多采用反应式状态到动作的映射,常导致短视行为,在动态场景中鲁棒性差。本文提出F1,一种预训练的VLA框架,将视觉前瞻生成融入决策流程。F1采用混合变压器架构,包含感知、前瞻生成和控制专用模块,实现理解、生成与动作的衔接。其核心是下一尺度预测机制,可合成目标条件下的视觉前瞻作为显式规划目标。通过预测可能的未来视觉状态,F1将动作生成重构为前瞻引导的逆动力学问题,使动作隐式达成视觉目标。为赋予F1鲁棒且可迁移的能力,我们在涵盖136种多样化任务、超33万条轨迹的大型数据集上提出三阶段训练方案,增强模块化推理能力,并赋予模型可转移的视觉前瞻能力,这对复杂动态环境至关重要。在真实任务与仿真基准上的大量评估表明,F1持续优于现有方法,在任务成功率和泛化能力上均有显著提升。

原文摘要 · Abstract (English)

Executing language-conditioned tasks in dynamic visual environments remains a central challenge in embodied AI. Existing Vision-Language-Action (VLA) models predominantly adopt reactive state-to-action mappings, often leading to short-sighted behaviors and poor robustness in dynamic scenes. In this paper, we introduce F1, a pretrained VLA framework which integrates the visual foresight generation into decision-making pipeline. F1 adopts a Mixture-of-Transformer architecture with dedicated modules for perception, foresight generation, and control, thereby bridging understanding, generation, and actions. At its core, F1 employs a next-scale prediction mechanism to synthesize goal-conditioned visual foresight as explicit planning targets. By forecasting plausible future visual states, F1 reformulates action generation as a foresight-guided inverse dynamics problem, enabling actions that implicitly achieve visual goals. To endow F1 with robust and generalizable capabilities, we propose a three-stage training recipe on an extensive dataset comprising over 330k trajectories across 136 diverse tasks. This training scheme enhances modular reasoning and equips the model with transferable visual foresight, which is critical for complex and dynamic environments. Extensive evaluations on real-world tasks and simulation benchmarks demonstrate F1 consistently outperforms existing approaches, achieving substantial gains in both task success rate and generalization ability.

具身AI视觉前瞻多模态决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。