构建多模态指代评估平台,验证模型在真实场景中的指物能力。
PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
- 设计三阶段评估体系:数据集、在线对抗测试、机器人实测
- Molmo-72B表现最优,微调后指物准确率显著提升
- 适合研究多模态交互、具身智能与人机协作的开发者
指物是将语言与视觉语境对齐的基本且直观机制,广泛应用于机器人、辅助技术及交互式AI系统。尽管近期多模态模型已具备指物能力,现有基准多局限于参照物定位任务。本文提出PointArena,一个涵盖多种推理场景的多模态指物评估平台。其包含三个部分:(1) Point-Bench,一个约1,000个指物任务的精选数据集,覆盖五类推理类别;(2) Point-Battle,一个基于网页的交互式竞技场,支持盲评配对模型比较,已收集超4,500次匿名投票;(3) Point-Act,一个真实世界的机器人操作系 统,允许用户直接在实际环境中评估模型指物能力。我们对主流开源与专有模型进行了全面评估,结果表明:Molmo-72B持续领先,专有模型性能也逐渐接近。此外,专门针对指物任务进行监督训练能显著提升模型表现。在多阶段评估中,各环节表现高度相关,凸显精确指物能力在连接抽象推理与具体现实动作中的关键作用。项目页:https://pointarena.github.io/
原文摘要 · Abstract (English)
Pointing serves as a fundamental and intuitive mechanism for grounding language within visual contexts, with applications spanning robotics, assistive technologies, and interactive AI systems. While recent multimodal models have started to support pointing capabilities, existing benchmarks typically focus only on referential object localization tasks. We introduce PointArena, a comprehensive platform for evaluating multimodal pointing across diverse reasoning scenarios. PointArena comprises three components: (1) Point-Bench, a curated dataset containing approximately 1,000 pointing tasks across five reasoning categories; (2) Point-Battle, an interactive, web-based arena facilitating blind, pairwise model comparisons, which has already gathered over 4,500 anonymized votes; and (3) Point-Act, a real-world robotic manipulation system allowing users to directly evaluate multimodal model pointing capabilities in practical settings. We conducted extensive evaluations of both state-of-the-art open-source and proprietary multimodal models. Results indicate that Molmo-72B consistently outperforms other models, though proprietary models increasingly demonstrate comparable performance. Additionally, we find that supervised training specifically targeting pointing tasks significantly enhances model performance. Across our multi-stage evaluation pipeline, we also observe strong correlations, underscoring the critical role of precise pointing capabilities in enabling multimodal models to effectively bridge abstract reasoning with concrete, real-world actions. Project page: https://pointarena.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。