开源多模态动作模型评估工具包,支持视觉、语言与动作协同测试。
An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
- 构建全开源的多模态动作模型评测框架与数据生态。
- 提供超1.3万亿标记的复合数据集,覆盖图像描述、机器人控制等任务。
- 适合研究通用智能体泛化能力与模型适配的开发者与学者使用。
近期多模态动作模型的发展为构建通用智能体系统提供了新方向,融合视觉理解、语言认知与动作生成能力。我们提出MultiNet——一个完全开源的基准测试套件及配套软件生态系统,用于严格评估和适配跨视觉、语言与动作领域的模型。建立了标准化的评估协议,用于评估视觉语言模型(VLMs)与视觉语言动作模型(VLAs),并开放获取相关数据、模型与评估工具。此外,提供一个包含超过1.3万亿标记的复合数据集,涵盖图像描述、视觉问答、常识推理、机器人控制、数字游戏、模拟运动/操作等多种任务。MultiNet基准、框架、工具包与评估工具已被应用于下游研究,揭示了VLA泛化能力的局限性。
原文摘要 · Abstract (English)
Recent innovations in multimodal action models represent a promising direction for developing general-purpose agentic systems, combining visual understanding, language comprehension, and action generation. We introduce MultiNet - a novel, fully open-source benchmark and surrounding software ecosystem designed to rigorously evaluate and adapt models across vision, language, and action domains. We establish standardized evaluation protocols for assessing vision-language models (VLMs) and vision-language-action models (VLAs), and provide open source software to download relevant data, models, and evaluations. Additionally, we provide a composite dataset with over 1.3 trillion tokens of image captioning, visual question answering, commonsense reasoning, robotic control, digital game-play, simulated locomotion/manipulation, and many more tasks. The MultiNet benchmark, framework, toolkit, and evaluation harness have been used in downstream research on the limitations of VLA generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。