首个用点光源动作评估多模态大模型的基准,发现模型理解动作能力严重不足。
Evaluating point-light biological motion in multimodal large language models
- 构建ActPLD基准,用人体关节点光动画测试模型动作理解能力。
- 各类模型在单人与互动场景下表现均差,动作与时空理解存在明显缺陷。
- 适合研究具身认知、动作理解与多模态模型评测的学者参考。
人类能从极简视觉线索中提取丰富语义信息,如点光源显示(PLDs),其仅由人体关键关节处的稀疏光点组成。这种能力在早期发育中即显现,主要源于人类的具身经验。由于PLDs将身体运动作为唯一意义来源,成为检验系统动作理解能力的关键刺激。本文提出ActPLD,是首个用于评估多模态大语言模型(MLLMs)在人体点光源动作理解上的基准。测试涵盖主流闭源与开源模型,覆盖单人及社会互动场景的PLDs。结果表明,所有模型表现均不理想,暴露出动作与时空理解方面的根本性差距。
原文摘要 · Abstract (English)
Humans can extract rich semantic information from minimal visual cues, as demonstrated by point-light displays (PLDs), which consist of sparse sets of dots localized to key joints of the human body. This ability emerges early in development and is largely attributed to human embodied experience. Since PLDs isolate body motion as the sole source of meaning, they represent key stimuli for testing the constraints of action understanding in these systems. Here we introduce ActPLD, the first benchmark to evaluate action processing in MLLMs from human PLDs. Tested models include state-of-the-art proprietary and open-source systems on single-actor and socially interacting PLDs. Our results reveal consistently low performance across models, introducing fundamental gaps in action and spatiotemporal understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。