arXiv:2507.09168cs.CV2025-07ICCV被引 4

通过单个分类器锚定提示,实现更稳定精准的文本引导图像3D编辑。

Stable Score Distillation

  • 用单一分类器和常数项空提示分支稳定优化过程。
  • 在2D/3D编辑中达到当前最优效果,收敛更快、结构更清晰。
  • 适合需要强风格转换与局部控制的生成任务开发者使用。

基于扩散模型的文本引导图像与3D编辑已取得进展,但如Delta Denoising Score等方法常面临稳定性差、空间控制弱及编辑强度不足的问题。根源在于依赖复杂的辅助结构,引入冲突优化信号,限制精确局部编辑。本文提出稳定分数蒸馏(SSD),通过将单一分类器锚定至源提示,提升编辑过程的稳定性与对齐性。具体地,SSD利用无分类器指导(CFG)方程实现跨提示对齐,并引入常数项空提示分支以稳定优化。该方法保持原始内容结构,确保编辑轨迹紧密贴合源提示,实现平滑、提示特异性的修改,同时维持周围区域连贯性。此外,引入提示增强分支提升风格转换等编辑强度。在NeRF及文本驱动风格编辑等2D/3D任务中,本方法达到当前最优性能,收敛更快、复杂度更低,为文本引导编辑提供高效鲁棒的解决方案。

原文摘要 · Abstract (English)

Text-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieves cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's structure and ensures that editing trajectories are closely aligned with the source prompt, enabling smooth, prompt-specific modifications while maintaining coherence in surrounding regions. Additionally, SSD incorporates a prompt enhancement branch to boost editing strength, particularly for style transformations. Our method achieves state-of-the-art results in 2D and 3D editing tasks, including NeRF and text-driven style edits, with faster convergence and reduced complexity, providing a robust and efficient solution for text-guided editing.

图像编辑扩散模型提示对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。