测试视觉模型能否像人一样看图主动决策
VisualActBench: Can VLMs See and Act like a Human?
- 设计新任务与基准,让模型仅凭图像自主推理行动
- 29个模型在1074段视频上测试,人类水平仍有差距
- 适合研究主动智能体和人机对齐的学者参考
视觉语言模型在感知和描述视觉环境方面取得显著进展,但其仅依赖视觉输入、无需文本提示即可主动推理并采取行动的能力仍鲜有研究。本文提出视觉动作推理新任务,并构建VisualActBench大规模基准,包含1,074段视频和3,733条人类标注动作,覆盖四个真实场景。每个动作标注了行动优先级(APL)和主动-被动类型,用于评估模型的人类对齐推理与价值敏感性。在该基准上评估29个VLM,发现前沿模型如GPT4o表现相对较强,但仍显著落后于人类,尤其在生成高优先级主动行为方面。结果揭示当前模型在理解复杂上下文、预判后果及契合人类决策框架方面的局限。VisualActBench为评估和提升主动视觉智能体的真实世界适用性奠定全面基础。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have achieved impressive progress in perceiving and describing visual environments. However, their ability to proactively reason and act based solely on visual inputs, without explicit textual prompts, remains underexplored. We introduce a new task, Visual Action Reasoning, and propose VisualActBench, a large-scale benchmark comprising 1,074 videos and 3,733 human-annotated actions across four real-world scenarios. Each action is labeled with an Action Prioritization Level (APL) and a proactive-reactive type to assess models' human-aligned reasoning and value sensitivity. We evaluate 29 VLMs on VisualActBench and find that while frontier models like GPT4o demonstrate relatively strong performance, a significant gap remains compared to human-level reasoning, particularly in generating proactive, high-priority actions. Our results highlight limitations in current VLMs' ability to interpret complex context, anticipate outcomes, and align with human decision-making frameworks. VisualActBench establishes a comprehensive foundation for assessing and improving the real-world readiness of proactive, vision-centric AI agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。