评测智能体执行精细动作的认知能力,发现主流大模型表现不佳。
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- 构建1368段视频+19562个问答对,覆盖4类认知能力。
- 主流多模态大模型在动作指令生成上准确率不足,高阶推理缺陷明显。
- 微调后模型在真实任务中性能显著提升,适合智能体研发者参考。
多模态大语言模型(MLLMs)在复杂物理环境中作为决策引擎展现出良好前景,但现有基准多关注高层规划或空间推理,忽视了实现物理交互所需的精细动作智能。为此,我们提出CFG-Bench,一个系统评估该能力的新基准。该基准包含1,368个精选视频与19,562个问答对,涵盖三个评估范式,聚焦四大认知能力:物理交互、时间因果关系、意图理解与评价判断。这些维度共同构成评估模型将视觉观察转化为可操作知识的能力框架,超越表层识别。我们在CFG-Bench上的全面评估显示,领先MLLMs在生成详细动作指令方面表现不佳,且在意图与评价等高阶推理上存在深层局限。此外,对数据进行监督微调(SFT)表明,直接训练模型表达精细动作可显著提升其在现有具身任务基准上的表现。分析揭示了当前瓶颈,并为开发更强大、更落地的具身智能体提供启示。项目页:https://cfg-bench.github.io/
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) show promising results as decision-making engines for embodied agents operating in complex, physical environments. However, existing benchmarks often prioritize high-level planning or spatial reasoning, leaving the fine-grained action intelligence required for embodied physical interaction underexplored. To address this gap, we introduce CFG-Bench, a new benchmark designed to systematically evaluate this crucial capability. CFG-Bench consists of 1,368 curated videos paired with 19,562 question-answer pairs spanning three evaluation paradigms targeting four cognitive abilities: 1) Physical Interaction, 2) Temporal-Causal Relation, 3) Intentional Understanding, and 4) Evaluative Judgment. Together, these dimensions provide a systematic framework for assessing a model's ability to translate visual observations into actionable knowledge, moving beyond mere surface-level recognition. Our comprehensive evaluation on CFG-Bench reveals that leading MLLMs struggle to produce detailed instructions for physical interactions and exhibit profound limitations in the higher-order reasoning of intention and evaluation. Moreover, supervised fine-tuning (SFT) on our data demonstrates that teaching an MLLMs to articulate fine-grained actions directly translates to significant performance gains on established embodied benchmarks. Our analysis highlights these limitations and offers insights for developing more capable and grounded embodied agents. Project page: https://cfg-bench.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。