无需训练即可精准控制图像中3D几何特征,提升设计效率。
GeoDiffusion: A Training-Free Framework for Accurate 3D Geometric Conditioning in Image Generation
- 用特定3D物体作为几何先验,定义关键点与空间关联。
- 在DragBench上实现更快更准的拖拽式编辑,精度显著提升。
- 适合工业设计、创意行业用户快速生成符合几何要求的图像。
在图像生成中实现精确的几何控制对工程与产品设计、创意产业至关重要,可准确调控图像空间中的3D物体特征。传统3D编辑方法耗时且需专业技能,而现有基于图像的生成方法在几何控制上准确性不足。为此,我们提出GeoDiffusion——一种无需训练的框架,用于图像生成中3D特征的精准高效几何调节。该框架采用类别特定的3D物体作为几何先验,定义3D空间中的关键点与参数化关联;通过渲染参考3D物体的视角一致性图像,并结合风格迁移以满足用户设定的外观要求。核心组件GeoDrag在DragBench基准上,显著提升基于几何引导与通用指令的拖拽编辑速度与精度。实验表明,GeoDiffusion可在多种迭代设计流程中实现精准几何修改。
原文摘要 · Abstract (English)
Precise geometric control in image generation is essential for engineering \& product design and creative industries to control 3D object features accurately in image space. Traditional 3D editing approaches are time-consuming and demand specialized skills, while current image-based generative methods lack accuracy in geometric conditioning. To address these challenges, we propose GeoDiffusion, a training-free framework for accurate and efficient geometric conditioning of 3D features in image generation. GeoDiffusion employs a class-specific 3D object as a geometric prior to define keypoints and parametric correlations in 3D space. We ensure viewpoint consistency through a rendered image of a reference 3D object, followed by style transfer to meet user-defined appearance specifications. At the core of our framework is GeoDrag, improving accuracy and speed of drag-based image editing on geometry guidance tasks and general instructions on DragBench. Our results demonstrate that GeoDiffusion enables precise geometric modifications across various iterative design workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。