构建首个聚焦非刚性运动的指令图像编辑基准,支持复杂动态变形。
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
- 基于扩散变换器设计新模型ByteMorpher,支持指令驱动编辑。
- 数据集含600万高分辨率图像对,覆盖多样非刚性运动类型。
- 适合研究视频生成、可控图像编辑与人形动作建模的学者。
以指令指导图像编辑来表现非刚性运动、相机视角变化、物体形变、人体关节动作及复杂交互,是计算机视觉中一个极具挑战但尚未充分探索的问题。现有方法和数据集多集中于静态场景或刚性变换,难以处理包含动态运动的表达性编辑。为填补这一空白,我们提出ByteMorph框架,包含大规模数据集ByteMorph-6M与基于扩散变压器(DiT)的强基线模型ByteMorpher。ByteMorph-6M包含超过600万张高分辨率图像编辑对,用于训练,并配有精心设计的评估基准ByteMorph-Bench。两者均涵盖多样化环境、人体与物体类别中的多种非刚性运动类型。数据集通过运动引导生成、分层合成技术与自动标注实现多样性、真实感与语义一致性。我们进一步对来自学术界与工业界的近期指令式图像编辑方法进行了全面评估。
原文摘要 · Abstract (English)
Editing images with instructions to reflect non-rigid motions, camera viewpoint shifts, object deformations, human articulations, and complex interactions, poses a challenging yet underexplored problem in computer vision. Existing approaches and datasets predominantly focus on static scenes or rigid transformations, limiting their capacity to handle expressive edits involving dynamic motion. To address this gap, we introduce ByteMorph, a comprehensive framework for instruction-based image editing with an emphasis on non-rigid motions. ByteMorph comprises a large-scale dataset, ByteMorph-6M, and a strong baseline model built upon the Diffusion Transformer (DiT), named ByteMorpher. ByteMorph-6M includes over 6 million high-resolution image editing pairs for training, along with a carefully curated evaluation benchmark ByteMorph-Bench. Both capture a wide variety of non-rigid motion types across diverse environments, human figures, and object categories. The dataset is constructed using motion-guided data generation, layered compositing techniques, and automated captioning to ensure diversity, realism, and semantic coherence. We further conduct a comprehensive evaluation of recent instruction-based image editing methods from both academic and commercial domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。