用生物运动动画测试大模型,发现多数模型连基本人形动作都做不好。
Can Large Models Fool the Eye? A New Turing Test for Biological Animation
- 通过点光源动画进行成对比较,让人类直观判断模型表现差异。
- 超4.5万次投票显示,90%以上模型无法生成自然的人类动作。
- 无需真实标签,适合评估模型在生物运动上的真实感知能力。
评估大模型能力并揭示其短板极具挑战。现有基准多依赖静态数据集的客观评分或模糊的文本对话偏好收集,难以提供即时、直观的性能反馈。本文提出生物运动竞技场(BioMotion Arena),一种基于视觉动画的大语言模型(LLMs)与多模态大语言模型(MLLMs)评估新框架。该方法借鉴生物体运动模式的视觉感知特性,利用点光源成像凸显模型间性能差异。我们采用成对比较方式,在90种生物运动变体上收集了超过4.5万次人类投票,涵盖53个主流模型。数据分析表明,众包人类评分与专家评分高度一致,验证了该框架的判别力。结果发现,包括顶尖开源模型InternVL3和专有模型Claude-4系列在内的90%以上模型,均无法生成基本的人形点光源序列,更无法实现平滑且符合生物学特征的运动。该框架可作为可视化性能评估的高难度基准,且不依赖真实标签。
原文摘要 · Abstract (English)
Evaluating the abilities of large models and manifesting their gaps are challenging. Current benchmarks adopt either ground-truth-based score-form evaluation on static datasets or indistinct textual chatbot-style human preferences collection, which may not provide users with immediate, intuitive, and perceptible feedback on performance differences. In this paper, we introduce BioMotion Arena, a novel framework for evaluating large language models (LLMs) and multimodal large language models (MLLMs) via visual animation. Our methodology draws inspiration from the inherent visual perception of motion patterns characteristic of living organisms that utilizes point-light source imaging to amplify the performance discrepancies between models. Specifically, we employ a pairwise comparison evaluation and collect more than 45k votes for 53 mainstream LLMs and MLLMs on 90 biological motion variants. Data analyses show that the crowd-sourced human votes are in good agreement with those of expert raters, demonstrating the superiority of our BioMotion Arena in offering discriminative feedback. We also find that over 90\% of evaluated models, including the cutting-edge open-source InternVL3 and proprietary Claude-4 series, fail to produce fundamental humanoid point-light groups, much less smooth and biologically plausible motions. This enables BioMotion Arena to serve as a challenging benchmark for performance visualization and a flexible evaluation framework without restrictions on ground-truth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。