arXiv:2601.18340cs.CV2026-01被引 1

评测视频生成模型对非刚性动态的编辑能力,揭示现有指标的不足。

Beyond Rigid: Benchmarking Non-Rigid Video Editing

  • 构建物理驱动的非刚性视频编辑基准,含180个视频与2340条指令。
  • 发现传统指标无法反映真实物理合理性,模型可能外观保真但动态失真。
  • 适合关注视频生成物理一致性与动态建模的研究者。

随着视频生成模型被期望操控物理动态,评估需超越外观保真度与语义对齐。非刚性视频编辑提供了一个独特检验场景,不同材质施加不同物理约束。本文提出NRVBench,一个用于非刚性视频编辑的诊断基准,任务是在保持无关区域不变的同时修改可变形运动,并维持材料特定的合理性。NRVBench包含六个物理基础类别、180个精心筛选的视频、2,340条细粒度编辑指令、360道多选题及像素级掩码。我们进一步提出NRVE-Acc,一种基于视觉语言模型的结构化评估协议,将编辑成功分解为指令遵循、材料感知变形合理性与运动线索的时间连贯性。在代表性推理时视频编辑方法上的实验显示,传统指标与物理感知编辑成功之间存在明显差距:即使保留外观或实现强全局对齐,模型在非刚性动态下仍可能失败。此外,我们引入VM-Edit,一种简单的区域条件编辑基线,仅释放前景并锁定背景,暴露出稳定与可塑性之间的权衡。

原文摘要 · Abstract (English)

As video generation models are increasingly expected to manipulate physical dynamics, there is a growing need to move evaluation beyond appearance fidelity and semantic alignment. Non-rigid video editing offers a uniquely revealing testbed, where distinct materials impose distinct physical constraints. In this paper, we introduce NRVBench, a diagnostic benchmark for non-rigid video editing, where the task is to modify deformable motion while preserving irrelevant regions and maintaining material-specific plausibility. NRVBench contains 180 curated videos across six physics-grounded categories, 2,340 fine-grained editing instructions, 360 multiple-choice questions, and pixel-accurate masks. We further propose NRVE-Acc, a structured VLM-based protocol that decomposes editing success into instruction following, material-aware deformation plausibility, and temporal coherence with motion cues. Experiments on representative inference-time video editing methods reveal a clear mismatch between conventional metrics and physics-aware perceptual editing success: methods that preserve appearance or achieve strong global alignment may still fail under non-rigid dynamics. We additionally introduce VM-Edit, a simple region-conditioned editing baseline that frees the foreground while locking the background, exposing the stability--plasticity trade-off.

视频编辑物理建模评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。