arXiv:2604.04911cs.CV2026-04被引 4

构建首个精细图像空间编辑基准,提升布局与视角控制精度。

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing

  • 提出SpatialEdit-Bench,联合评估视觉真实感与几何保真度。
  • 构建50万张合成数据集,支持物体与相机双视角精准变换。
  • 推出160亿参数模型,空间操作性能显著超越现有方法。

图像空间编辑通过几何驱动实现对物体布局和摄像机视角的精确控制。当前模型在细粒度空间操作上表现不足,亟需专门评估体系。本文贡献如下:(i) 提出SpatialEdit-Bench,一个完整基准,通过视角重建与构图分析联合评估视觉合理性和几何保真度;(ii) 为解决可扩展训练的数据瓶颈,构建SpatialEdit-500k,基于可控Blender流水线生成包含多样化背景与系统性相机轨迹的合成图像,提供物体与相机为中心操作的精确真实标签;(iii) 基于此数据,开发SpatialEdit-16B基线模型,在通用编辑任务中表现良好,且在空间操作任务上显著优于先前方法。所有资源将公开于https://github.com/EasonXiao-888/SpatialEdit。

原文摘要 · Abstract (English)

Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained spatial manipulations, motivating a dedicated assessment suite. Our contributions are listed: (i) We introduce SpatialEdit-Bench, a complete benchmark that evaluates spatial editing by jointly measuring perceptual plausibility and geometric fidelity via viewpoint reconstruction and framing analysis. (ii) To address the data bottleneck for scalable training, we construct SpatialEdit-500k, a synthetic dataset generated with a controllable Blender pipeline that renders objects across diverse backgrounds and systematic camera trajectories, providing precise ground-truth transformations for both object- and camera-centric operations. (iii) Building on this data, we develop SpatialEdit-16B, a baseline model for fine-grained spatial editing. Our method achieves competitive performance on general editing while substantially outperforming prior methods on spatial manipulation tasks. All resources will be made public at https://github.com/EasonXiao-888/SpatialEdit.

图像编辑空间操控合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。