为视频生成模型的刚体物理行为设计了可量化评估基准
RigidBench: Evaluating Rigid-Body Physics in Video Generation Models

- 基于模拟器构建五类刚体任务,分离评估运动、几何、身份等维度
- 8个模型在100例测试中表现差异大,3D轨迹误差与图像相似度负相关
- 提供5000段带精确状态的训练视频,可用于模型微调与机制分析
视频生成模型日益用于预测场景中的未来变化,但常用评价指标难以判断生成物体是否正确运动。运动、几何、身份、背景稳定性与视觉相似性可独立出错,而整体评分常将这些错误混合。我们提出RigidBench,一个基于模拟器的基准,将生成结果与同一初始帧和运动描述下的参考轨迹进行对比。该基准包含五类刚体任务,涵盖不同物体、材料、视角及室内外场景,提供每帧掩码、深度图、6-DoF轨迹和接触信息用于评分。我们在相同100个样本上对8个模型进行十项独立测量,结果表明排名高度依赖于评估维度:无模型在所有指标上领先;模型平均表现中,SSIM越高,3D轨迹误差越大(r = 0.89)。RigidBench还包含5,000段带精确模拟状态的训练视频,用于微调和分析Wan 2.2 TI2V-5B。全量微调使3D轨迹误差降低约20%,而SSIM几乎不变;教师强制探测与定向干预显示,物体位置在整个扩散变换器中被持续表征并参与去噪计算。
原文摘要 · Abstract (English)
Video models are increasingly used to predict what happens next in a scene, yet the metrics commonly used to compare their outputs say little about whether the predicted objects move correctly. Motion, geometry, identity, background stability, and visual similarity can fail independently, but whole-frame scores often mix these errors together. We introduce RigidBench, a simulator-grounded benchmark that compares a generated continuation with a reference rollout from the same initial frame and motion description. Its five rigid-body tasks vary objects, materials, viewpoints, and indoor and outdoor scenes, with per-frame masks, depth, 6-DoF trajectories, and contacts available for scoring. We evaluate eight models on the same 100 examples with ten measurements that keep these aspects separate. The resulting rankings depend strongly on what is measured: no model leads on all ten, and across model means, higher SSIM accompanies larger 3D trajectory error (r = 0.89). RigidBench also includes 5,000 training videos with exact simulator state, which we use to fine-tune and analyze Wan 2.2 TI2V-5B. Full fine-tuning reduces 3D trajectory error by about 20% with almost no change in SSIM, while teacher-forced probes and targeted interventions show that object position is represented throughout Wan's diffusion transformer and used by its denoising computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。