arXiv:2605.14843cs.CV2026-05

评测视频生成模型的机械运动一致性,发现现有模型常违背物理约束。

MechVerse: Evaluating Physical Motion Consistency in Video Generation Models

论文配图:MechVerse: Evaluating Physical Motion Consistency in Video Generation Models
图 1 · 摘自论文原文
  • 构建包含141类机械结构的合成视频基准集,分三阶递增复杂度
  • 模型在保持画面流畅时仍频繁违反刚体、连接与传动规则,错误随耦合复杂度上升
  • 适合关注物理真实性视频生成的研究者和开发者

文本和图像条件下的视频生成模型虽具备强视觉保真度和时间连贯性,但常无法生成符合运动学与几何约束的运动。在这些场景中,物体部件必须保持刚性、维持接触或耦合关系,并沿连接部件一致传递运动。这一要求在可动机械组件中尤为明显,其运动受刚性链接几何、接触/耦合关系及运动链传输约束。生成视频可能看似合理,却违背预期机制,如旋转应平移的部件、变形刚性构件、破坏部件间耦合或未带动下游组件。为此,我们引入MechVerse,一个用于评估图像到视频生成机械一致性的基准。MechVerse包含1,357个机械组件的21,156段合成视频,覆盖141个类别,分为三阶递增的运动学复杂度:独立运动、成对耦合、密集耦合多部件机构。每段视频配有结构化提示,描述部件身份、固定支撑、运动组件、运动基元、方向、速度/范围及部件间依赖关系。我们使用标准视频指标、指令遵循得分及人工判断运动正确性与运动耦合性,评估专有、开源和微调的图像到视频模型。结果表明,当前模型可在保持外观与平滑性的同时,生成不符合机械可行性的运动,且错误随耦合复杂度增加而加剧。MechVerse为衡量和改进基于图像与语言输入的机制感知视频生成提供了基准。

原文摘要 · Abstract (English)

Text- and image-conditioned video generation models have achieved strong visual fidelity and temporal coherence, but they often fail to generate motion governed by kinematic and geometric constraints. In these settings, object parts must remain rigid, maintain contact or coupling with neighboring components, and transfer motion consistently across connected parts. These requirements are especially explicit in articulated mechanical assemblies, where motion is constrained by rigid-link geometry, contact/coupling relations, and transmission through kinematic chains. A generated video may therefore appear plausible while violating the intended mechanism, such as rotating a part that should translate, deforming a rigid component, breaking coupling between parts, or failing to move downstream components. To evaluate this gap, We introduce MechVerse, a benchmark for mechanically consistent image-to-video generation. MechVerse contains 21,156 synthetic clips from 1,357 mechanical assemblies across 141 categories, organized into three tiers of increasing kinematic complexity: independent articulation, pairwise coupling, and densely coupled multi-part mechanisms. Each clip is paired with a structured prompt describing part identities, stationary supports, moving components, motion primitives, direction, speed/extent, and inter-part dependencies. We evaluate proprietary, open-source, and fine-tuned image-to-video models using standard video metrics, instruction-following scores, and human judgments of motion correctness and kinematic coupling. Results show that current models can preserve appearance and smoothness while failing to generate mechanically admissible motion, with errors increasing as coupling complexity grows. MechVerse provides a benchmark for measuring and improving mechanism-aware video generation from image and language inputs.

视频生成物理一致性机械运动基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。