arXiv:2506.00893cs.CLcs.AI2025-06被引 6

评测大模型对物体使用可能性的理解能力,发现性能远低于人类。

Affordance Benchmark for MLLMs

  • 构建双维度基准A4Bench,涵盖固有属性与动态情境下的使用可能性
  • 顶级模型准确率仅18.05%,人类最佳达85.34%,差距显著
  • 揭示大模型在情境理解上的短板,适合关注AI交互安全的研究者

具身理论认为环境本身蕴含可行动作的可能,影响感知与行为。尽管多模态大语言模型(MLLM)在视觉-语言任务中表现优异,其对具身性(affordance)的感知能力——这对直觉且安全的交互至关重要——仍鲜被研究。为此,我们提出A4Bench,一个新型基准,从两个维度评估MLLM的具身感知能力:1)构成性具身,通过1282个跨九个子领域的问答对评估对象固有属性理解;2)转化性具身,通过718个具有挑战性的问答对考察动态、上下文相关的复杂具身(如误导性、时间依赖性、文化或个体差异)。我们评估了17个MLLM(9个专有模型和8个开源模型),并与人类表现对比。结果显示,专有模型普遍优于开源模型,但所有模型均远低于人类,尤其在转化性具身上表现更差。即使顶尖模型Gemini-2.0-Pro的总体精确匹配准确率也仅为18.05%,而人类表现最佳达85.34%,最差为81.25%。这些发现凸显了当前MLLM在环境理解上的关键缺陷,为构建更具鲁棒性、上下文敏感的AI系统提供了基础。

原文摘要 · Abstract (English)

Affordance theory suggests that environments inherently provide action possibilities shaping perception and behavior. While Multimodal Large Language Models (MLLMs) achieve strong performance in vision-language tasks, their ability to perceive affordance, which is crucial for intuitive and safe interactions, remains underexplored. To address this, we introduce **A4Bench**, a novel benchmark designed to evaluate the affordance perception abilities of MLLMs across two dimensions: 1) Constitutive Affordance, assessing understanding of inherent object properties through 1,282 questionanswer pairs spanning nine sub-disciplines, and 2) Transformative Affordance, probing dynamic and contextual nuances (e.g., misleading, time-dependent, cultural, or individual-specific affordance) with 718 challenging question-answer pairs. We evaluate 17 MLLMs (nine proprietary and eight open-source) and compare them to human performance. Results show that proprietary models generally outperform open-source ones, yet all models perform far below humans, especially in transformative affordance. Furthermore, even top-performing models, such as Gemini-2.0-Pro (18.05% overall exact match accuracy), significantly lag behind human performance (best: 85.34%, worst: 81.25%). These findings highlight critical gaps in environmental understanding of MLLMs and provide a foundation for advancing AI systems toward more robust, context-aware interactions.

多模态具身认知评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。