arXiv:2607.22393cs.AIcs.CV2026-07

评测视觉语言模型在3D场景中的综合行动能力。

SceneActBench: Can Agents Act on the 3D Scenes They See?

论文配图:SceneActBench: Can Agents Act on the 3D Scenes They See?
图 1 · 摘自论文原文
  • 构建统一环境下的五类3D任务评估框架
  • 11种模型平均得分38.6-50.2,无模型全任务表现优
  • 揭示模型在多物体交互中的失败模式,适合研究者参考

视觉语言模型(VLM)代理正从仅描述3D场景转向实际操作。现有3D基准多评价文本回复或单对象操作,未能全面评估代理在完整多物体3D场景中的行动能力。本文提出SceneActBench,一个基于统一代理-环境循环的基准,涵盖五类视觉条件下的3D动作任务。给定PNG图像或采样视频帧,并在适用时提供3D资产,代理在3D环境中执行动作。每个最终输出通过任务特定几何指标与隐藏真实值对比评估。该基准包含210个原始实例,生成520个任务案例,涵盖配对输入条件。所有任务均通过单一固定代理流程运行,确保公平比较。在11种专有VLM配置中,整体得分范围为38.6-50.2,且无一模型在所有任务上表现稳定。进一步分析揭示了失败发生的位置与方式。

原文摘要 · Abstract (English)

Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D environment. We evaluate each final output against hidden ground truth with task-specific geometric metrics. SceneActBench comprises five tasks built from 210 source instances, yielding 520 task cases including paired input conditions. Every task runs through one fixed agent loop to keep the comparison fair. Across eleven proprietary VLM configurations, Overall scores span 38.6-50.2, and none performs consistently well across tasks. We further analyse where and how failures manifest.

3D视觉智能体基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。