构建物理AI统一评测基准,揭示现有模型在真实动态预测上的短板。
PAI-Bench: A Comprehensive Benchmark For Physical AI
- 设计2808个真实场景的多任务评测,涵盖视频生成与理解
- 发现生成模型虽视觉逼真但物理逻辑常出错,大模型预测能力弱
- 适合研究物理感知、具身智能和多模态生成的学者参考
物理AI旨在发展能感知和预测现实世界动态的模型,但当前多模态大语言模型与视频生成模型在支持此类能力方面的程度尚不明确。我们提出物理AI基准(PAI-Bench),一个统一且全面的评测体系,用于评估视频生成、条件视频生成和视频理解中的感知与预测能力。该基准包含2,808个真实世界案例,配备任务对齐的度量指标,以捕捉物理合理性与领域特定推理。系统评估显示,尽管视频生成模型具有强视觉保真度,但常难以维持物理上连贯的动力学;多模态大语言模型在预测和因果解释方面表现有限。这表明当前系统在应对物理AI的感知与预测需求方面仍处于初级阶段。总体而言,PAI-Bench为物理AI的评估提供了真实基础,并指出了未来系统需弥补的关键差距。
原文摘要 · Abstract (English)
Physical AI aims to develop models that can perceive and predict real-world dynamics; yet, the extent to which current multi-modal large language models and video generative models support these abilities is insufficiently understood. We introduce Physical AI Bench (PAI-Bench), a unified and comprehensive benchmark that evaluates perception and prediction capabilities across video generation, conditional video generation, and video understanding, comprising 2,808 real-world cases with task-aligned metrics designed to capture physical plausibility and domain-specific reasoning. Our study provides a systematic assessment of recent models and shows that video generative models, despite strong visual fidelity, often struggle to maintain physically coherent dynamics, while multi-modal large language models exhibit limited performance in forecasting and causal interpretation. These observations suggest that current systems are still at an early stage in handling the perceptual and predictive demands of Physical AI. In summary, PAI-Bench establishes a realistic foundation for evaluating Physical AI and highlights key gaps that future systems must address.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。