arXiv:2512.11315cs.LG2025-12

测试视觉语言动作模型跨领域泛化能力,发现现有模型表现不稳定。

Benchmarking the Generality of Vision-Language-Action Models

  • 构建统一基准MultiNet v1.0,覆盖6类核心能力域。
  • 三款主流模型在新任务上性能普遍下降,最高降幅达45%。
  • 适合关注通用智能评估与模型鲁棒性的研究者参考。

通用多模态智能体应整合感知、语言与控制,在多样化真实场景中稳健运行。然而当前评估体系分散于孤立基准,难以判断基础模型是否真正超越训练分布。本文提出MultiNet v1.0,一个统一基准,用于衡量视觉语言模型(VLMs)和视觉语言动作模型(VLAs)在六大基础能力域——视觉定位、空间推理、工具使用、物理常识、多智能体协同与连续机器人控制——上的跨领域泛化能力。评估GPT-5、Pi0与Magma后发现,无一模型展现出持续泛化能力。所有模型在未见领域、陌生模态或跨领域任务迁移时均出现显著性能下降,最高降幅达45%。失败表现包括模态错位、输出格式不稳与灾难性知识退化。结果揭示了通用智能愿景与当前基础模型实际能力之间的持续鸿沟。MultiNet v1.0为诊断该差距提供了标准化评估基底,代码、数据与排行榜已公开。

原文摘要 · Abstract (English)

Generalist multimodal agents are expected to unify perception, language, and control - operating robustly across diverse real world domains. However, current evaluation practices remain fragmented across isolated benchmarks, making it difficult to assess whether today's foundation models truly generalize beyond their training distributions. We introduce MultiNet v1.0, a unified benchmark for measuring the cross domain generality of vision language models (VLMs) and vision language action models (VLAs) across six foundational capability regimes. Visual grounding, spatial reasoning, tool use, physical commonsense, multi agent coordination, and continuous robot control. Evaluating GPT 5, Pi0, and Magma, we find that no model demonstrates consistent generality. All exhibit substantial degradation on unseen domains, unfamiliar modalities, or cross domain task shifts despite strong performance within their training distributions.These failures manifest as modality misalignment, output format instability, and catastrophic knowledge degradation under domain transfer.Our findings reveal a persistent gap between the aspiration of generalist intelligence and the actual capabilities of current foundation models.MultiNet v1.0 provides a standardized evaluation substrate for diagnosing these gaps and guiding the development of future generalist agents.Code, data, and leaderboards are publicly available.

多模态通用智能评估基准机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。