arXiv:2606.15869cs.CV2026-06被引 5

提出高效通用的自动驾驶世界动作模型,推理速度更快且泛化更强。

Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation

论文配图:Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation
图 1 · 摘自论文原文
  • 分离视频生成与动作预测模块,用混合专家架构分别建模
  • 在多个基准上达到顶尖性能,推理延迟显著降低
  • 适合需要快速响应的机器人导航和自动驾驶场景

世界动作模型(WAMs)在自动驾驶和城市导航中展现出巨大潜力。现有方法基于视觉-语言-动作模型或视频生成模型,存在两大缺陷:(1) 推理时需预测未来观测导致延迟高;(2) 视频与动作建模紧密耦合,引发表征错配、泛化能力下降。为此,我们提出Metis,一种端到端的解耦式WAM框架。Metis采用混合注意力变压器架构,为视频生成和动作预测分别设置专用专家,保持各任务的分布特性。为提升效率,引入非对称注意力掩码,在联合训练中允许动作模型在推理时跳过显式视频生成,实现训练-推理一致性,大幅降低计算开销且不影响规划性能。大量实验表明,Metis在NAVSIM navhard/navtest及CityWalker导航基准上均达领先水平,验证了其跨任务的泛化性与高效性。真实机器人部署进一步证实该方法的实际可行性。

原文摘要 · Abstract (English)

World action models~(WAMs) have shown great promise for autonomous driving and urban navigation. Built upon Vision-Language-Action models or video generation models, existing approaches suffer key limitations: (1) High inference latency due to future observation prediction at test time, and (2) tightly coupled video and action modeling leading to representational mismatch and degraded generalization. To address both issues, we propose Metis, an end-to-end WAM framework that decouples video generation and action prediction. Specifically, Metis employs a Mixture-of-Transformers architecture with dedicated experts for video generation and action prediction, preserving the intrinsic distributional properties of each task. To enhance efficiency, we introduce an asymmetric attention mask that enables joint training of both experts while allowing the action model to bypass explicit video generation during inference. This design ensures training-inference consistency and significantly reduces computational costs without compromising planning performance. Extensive experiments demonstrate state-of-the-art performance on the NAVSIM navhard and navtest benchmarks and the CityWalker navigation benchmark, validating both the generalizability and efficiency across diverse tasks. Real-robot deployments further confirm the practical feasibility of our approach.

自动驾驶世界模型高效推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。