arXiv:2606.01955cs.ROcs.CV2026-06被引 9

用语义事件重构视频动作模型,实现更自然的跨场景通用推理。

WALL-WM: Carving World Action Modeling at the Event Joints

论文配图:WALL-WM: Carving World Action Modeling at the Event Joints
图 1 · 摘自论文原文
  • 以语义事件为学习单元,替代固定时长动作片段
  • 在真实世界多场景下实现领先性能,支持可变长度执行
  • 适合需要跨任务泛化的机器人控制与智能体系统

WALL-WM 是一种世界动作模型(World Action Model),将视频动作学习从以片段为中心的优化转向以事件为基础的视觉-语言-动作预训练,以语义连贯的动作事件作为学习的基本单元。现有 WAM 通常基于多模态或视频基础模型初始化,并直接根据当前观测和指令优化固定长度的动作片段。这种片段中心范式存在根本性粒度不匹配:语言描述的是语义目标与事件,视觉呈现连续场景动态,动作则在控制级时间尺度上运行。强制三者进入同一固定长度预测窗口,使 VLA 训练退化为短时相关性拟合。WALL-WM 通过围绕语义事件组织监督信号与数据,采用事件级标注与聚类平衡采样构建数据生态,实现对多样化行为、场景与任务结构的可扩展学习。基于同一事件预训练主干网络,其支持两种互补推理模式:事件模式接收下一事件描述,支持可变长度执行块;统一模式使用带阶梯解码的视觉语言模型,条件化传统固定长度片段推理,同时保持梯度连续的 VLA 路径。结合基于 Muon 优化器的大规模预训练基础设施,WALL-WM 提供了通用型 WAM 的实用规模化方案。实验表明,WALL-WM 在语言、场景与任务间具有广泛泛化能力,在大规模真实世界泛化评估中达到当前最优表现。

原文摘要 · Abstract (English)

WALL-WM is a World Action Model that shifts video-action learning from chunk-centric optimization to event-grounded Vision-Language-Action pretraining, using semantically coherent action events as the atomic unit of learning. Existing WAMs commonly initialize from multimodal or video foundation models and then optimize fixed-length action chunks conditioned directly on the current observation and instruction. Although convenient, this chunk-centric formulation creates a fundamental granularity mismatch. Language describes semantic goals and events, vision evolves through continuous scene dynamics, and actions operate at control-level timescales; forcing all three into the same fixed-length prediction window turns VLA training into short-horizon correlation fitting. WALL-WM addresses this mismatch by organizing both supervision and data around semantic events. Specifically, it pairs event-grounded VLA pretraining with a data ecosystem built from event-level captions and cluster-balanced sampling, enabling scalable learning over diverse behaviors, scenes, and task structures. From the same event-pretrained backbone, WALL-WM supports two complementary inference modes. The event mode consumes next-event descriptions and enables variable-length execution chunks, while the unified mode uses a VLM with Staircase Decoding to condition conventional fixed-length chunk inference while preserving a gradient-continuous VLA path. Together with Muon-optimizer-based large-scale pretraining infrastructure, WALL-WM provides a practical scale-up recipe for general-purpose WAMs. Experiments show that WALL-WM generalizes broadly across language, scenes, and tasks, achieving state-of-the-art performance in large-scale real-world generalization evaluation.

动作建模视觉语言机器人控制事件感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。