用文字直接控制图像生成3D模型,又快又准。
Feedforward 3D Editing via Text-Steerable Image-to-3D
- 借鉴ControlNet思路,通过前向传播实现文本控制3D生成
- 在10万数据上训练,指令遵循更准确,原图一致性更好
- 比现有方法快2.4到28.5倍,适合快速设计与交互应用
图像到3D生成的进展为设计、AR/VR和机器人等领域带来巨大潜力。然而,要将AI生成的3D资产用于实际应用,关键在于能够轻松编辑。本文提出一种前馈式方法Steer3D,为图像到3D模型增加文本可操控性,实现语言驱动的3D资产编辑。该方法受ControlNet启发,通过适配图像到3D生成任务,可在一次前向传播中直接实现文本控制。我们构建了可扩展的数据生成引擎,并采用基于流匹配训练与直接偏好优化(DPO)的两阶段训练方案。相比现有方法,Steer3D在遵循语言指令方面表现更佳,且与原始3D资产的一致性更高,同时速度提升2.4倍至28.5倍。实验表明,在仅使用10万条数据的情况下,即可实现对预训练图像到3D生成模型的新模态(文本)控制。项目主页:https://glab-caltech.github.io/steer3d/
原文摘要 · Abstract (English)
Recent progress in image-to-3D has opened up immense possibilities for design, AR/VR, and robotics. However, to use AI-generated 3D assets in real applications, a critical requirement is the capability to edit them easily. We present a feedforward method, Steer3D, to add text steerability to image-to-3D models, which enables editing of generated 3D assets with language. Our approach is inspired by ControlNet, which we adapt to image-to-3D generation to enable text steering directly in a forward pass. We build a scalable data engine for automatic data generation, and develop a two-stage training recipe based on flow-matching training and Direct Preference Optimization (DPO). Compared to competing methods, Steer3D more faithfully follows the language instruction and maintains better consistency with the original 3D asset, while being 2.4x to 28.5x faster. Steer3D demonstrates that it is possible to add a new modality (text) to steer the generation of pretrained image-to-3D generative models with 100k data. Project website: https://glab-caltech.github.io/steer3d/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。