arXiv:2605.27351cs.CV2026-05被引 6

通过语义部件变换训练3D编辑模型,实现高质量可控编辑。

Feedforward 3D Editing Learns from Semantic-Part Transformation

论文配图:Feedforward 3D Editing Learns from Semantic-Part Transformation
图 1 · 摘自论文原文
  • 基于语义3D部件设计编辑数据集,提升定位精度
  • 提出PartFlow模型,实现无掩码推理的高保真编辑
  • 适合需要精细3D内容生成的研究者与开发者

3D编辑是可扩展3D内容创作的核心能力。尽管图像编辑已迈向大规模前馈生成范式,3D AI生成仍以无需训练的编辑流程为主。前馈3D编辑的核心挑战在于缺乏高质量成对监督。可编辑3D资产需同时保持几何结构、多视角一致性、结构连贯性及局部编辑可控性。现有3D编辑数据集多依赖独立生成资产、图像媒介重建或窄编辑分类,导致定位不准、保留弱、编辑边界模糊和语义一致性差。本文提出新视角:可扩展的前馈3D编辑应从语义部件变换中学习。基于此,我们构建了Pxform数据集,包含超过10万对跨七类编辑的一致前后编辑样本。我们的方法将编辑直接锚定在语义3D部件上。在此基础上,提出PartFlow模型,将源感知潜在控制注入预训练3D生成先验。PartFlow引入掩码感知速度保留与渲染空间一致性监督,联合提升编辑保真度与源保留效果,且推理时无需3D编辑掩码。大量实验表明,高质量语义部件监督显著提升可扩展3D编辑性能,使PartFlow在几何与外观编辑基准上达到领先水平。

原文摘要 · Abstract (English)

3D editing is a fundamental capability for scalable 3D content creation. While image editing has rapidly evolved toward large-scale feedforward generative paradigms, 3D AI generation remains dominated by training-free editing pipelines. A central challenge of feedforward 3D editing lies in the lack of high-quality paired supervision. Editable 3D assets require simultaneous preservation of geometry, multi-view consistency, structural coherence, and localized edit controllability. Existing 3D editing datasets often rely on independently generated assets, image-mediated reconstruction or narrow edit taxonomies, leading to inaccurate localization, weak preservation, blurred edit boundaries, and limited semantic consistency. In this work, we introduce a new perspective: scalable feedforward 3D editing should be learned from semantic-part transformations. Based on this insight, we propose Pxform, a high-quality 3D editing dataset with over 100K consistent before/after editing pairs across seven edit types. Instead of treating objects as unstructured shapes, our pipeline grounds edits directly in semantic 3D parts. Built upon Pxform, we further propose PartFlow, a feedforward 3D editing network that injects source-aware latent control into pretrained 3D generative priors. PartFlow introduces mask-aware velocity preservation and render-space consistency supervision to jointly improve edit fidelity and source preservation, while requiring no 3D edit mask during inference. Extensive experiments demonstrate that high-quality semantic-part supervision substantially improves scalable 3D editing, enabling PartFlow to achieve state-of-the-art performance on both geometric and appearance editing benchmarks.

3D编辑语义分割生成模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。