用少量图像实现无需微调的3D一致编辑,靠的是重用扩散模型的隐含3D感知。
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
- 复用预训练扩散模型的3D先验,避免每场景单独优化。
- 仅需1-2张图即可生成多视角一致的编辑结果,且效果优于现有方法。
- 适合想快速生成高质量3D内容的研究者和创作者使用。
我们提出Tinker,一个无需场景微调即可在单次或少数几次输入下实现高保真3D编辑的通用框架。与以往需大量场景优化以保证多视角一致性的方法不同,Tinker仅需1到2张图像即可生成多视角一致的编辑结果。其核心在于重用预训练扩散模型的潜在3D感知能力。为推动该领域研究,我们构建了首个大规模多视角编辑数据集及数据流水线,覆盖多样场景与风格。基于此,我们开发了两个新组件:(1) 参考式多视角编辑器,实现跨视角一致的精准编辑;(2) 任意视角到视频合成器,利用视频扩散模型的空间-时间先验,从稀疏输入中完成高质量场景补全与新视角生成。大量实验表明,Tinker在编辑、新视角生成和渲染增强任务上均达到当前最优性能,显著降低了通用3D内容创作的门槛,代表向真正可扩展的零样本3D编辑迈出关键一步。
原文摘要 · Abstract (English)
We introduce Tinker, a versatile framework for high-fidelity 3D editing that operates in both one-shot and few-shot regimes without any per-scene finetuning. Unlike prior techniques that demand extensive per-scene optimization to ensure multi-view consistency or to produce dozens of consistent edited input views, Tinker delivers robust, multi-view consistent edits from as few as one or two images. This capability stems from repurposing pretrained diffusion models, which unlocks their latent 3D awareness. To drive research in this space, we curate the first large-scale multi-view editing dataset and data pipeline, spanning diverse scenes and styles. Building on this dataset, we develop our framework capable of generating multi-view consistent edited views without per-scene training, which consists of two novel components: (1) Referring multi-view editor: Enables precise, reference-driven edits that remain coherent across all viewpoints. (2) Any-view-to-video synthesizer: Leverages spatial-temporal priors from video diffusion to perform high-quality scene completion and novel-view generation even from sparse inputs. Through extensive experiments, Tinker significantly reduces the barrier to generalizable 3D content creation, achieving state-of-the-art performance on editing, novel-view synthesis, and rendering enhancement tasks. We believe that Tinker represents a key step towards truly scalable, zero-shot 3D editing. Project webpage: https://aim-uofa.github.io/Tinker
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。