构建首个聚焦动作编辑的高质量数据集,提升视频生成中动作连贯性。
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
- 提出运动中心图像编辑任务,强调动作与结构一致性。
- 基准测试显示现有扩散模型在动作编辑上仍表现不佳。
- 设计MotionNFT框架,通过运动流匹配提升编辑精度。
我们提出MotionEdit,一个面向运动中心图像编辑的新数据集,即在保持主体身份、结构和物理合理性的同时修改动作与交互。与以往侧重静态外观变化或仅含稀疏低质运动编辑的数据集不同,MotionEdit提供从连续视频中提取并验证的高保真图像对,真实呈现运动变换。该任务兼具科学挑战性与实际意义,可支持帧控视频合成与动画生成等应用。为此,我们引入MotionEdit-Bench基准,采用生成式、判别式和偏好评估等多种指标衡量模型性能。结果表明,现有主流基于扩散的编辑模型在该任务上仍面临巨大挑战。为填补这一差距,我们提出MotionNFT(运动引导的负向感知微调)框架,通过计算输入与模型输出间运动流与真实运动的匹配度,生成运动对齐奖励,引导模型实现更准确的动作转换。在FLUX.1 Kontext和Qwen-Image-Edit上的实验表明,MotionNFT在不牺牲通用编辑能力的前提下,持续提升基础模型的编辑质量与运动保真度,验证了其有效性。
原文摘要 · Abstract (English)
We introduce MotionEdit, a novel dataset for motion-centric image editing-the task of modifying subject actions and interactions while preserving identity, structure, and physical plausibility. Unlike existing image editing datasets that focus on static appearance changes or contain only sparse, low-quality motion edits, MotionEdit provides high-fidelity image pairs depicting realistic motion transformations extracted and verified from continuous videos. This new task is not only scientifically challenging but also practically significant, powering downstream applications such as frame-controlled video synthesis and animation. To evaluate model performance on the novel task, we introduce MotionEdit-Bench, a benchmark that challenges models on motion-centric edits and measures model performance with generative, discriminative, and preference-based metrics. Benchmark results reveal that motion editing remains highly challenging for existing state-of-the-art diffusion-based editing models. To address this gap, we propose MotionNFT (Motion-guided Negative-aware Fine Tuning), a post-training framework that computes motion alignment rewards based on how well the motion flow between input and model-edited images matches the ground-truth motion, guiding models toward accurate motion transformations. Extensive experiments on FLUX.1 Kontext and Qwen-Image-Edit show that MotionNFT consistently improves editing quality and motion fidelity of both base models on the motion editing task without sacrificing general editing ability, demonstrating its effectiveness. Our code is at https://github.com/elainew728/motion-edit/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。