arXiv:2511.11692cs.LGcs.AI2025-11AAAI被引 2

让3D生成更连贯:用动态图像锚点稳定文本到3D的生成过程

AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation

  • 引入动态图像锚点,使生成过程随中间图像变化而调整
  • 在复杂提示下显著提升语义一致性与细节质量
  • 轻量级设计,保持高效且无需额外训练开销

基于优化的文本到3D方法通过得分蒸馏采样(SDS)从2D生成模型中提取指导,但隐式地将这些指导视为静态。本文指出,忽略源动态会导致不一致轨迹,抑制或融合语义线索,引发“语义过平滑”伪影。为此,我们将文本到3D优化重新建模为将动态演化源分布映射到固定目标分布的问题。在双条件潜在空间中,同时依赖文本提示和中间渲染图像进行条件控制,发现图像条件可自然锚定当前源分布。基于此,提出AnchorDS,一种提供状态锚定指导的改进得分蒸馏机制,有效稳定生成过程。进一步设计轻量级滤波与微调策略,惩罚错误源估计,以极低开销精炼锚点。实验表明,AnchorDS在复杂提示下生成更精细的细节、更自然的色彩,并显著增强语义一致性,同时保持高效率,优于现有方法。

原文摘要 · Abstract (English)

Optimization-based text-to-3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent trajectories that suppress or merge semantic cues, leading to "semantic over-smoothing" artifacts. As such, we reformulate text-to-3D optimization as mapping a dynamically evolving source distribution to a fixed target distribution. We cast the problem into a dual-conditioned latent space, conditioned on both the text prompt and the intermediately rendered image. Given this joint setup, we observe that the image condition naturally anchors the current source distribution. Building on this insight, we introduce AnchorDS, an improved score distillation mechanism that provides state-anchored guidance with image conditions and stabilizes generation. We further penalize erroneous source estimates and design a lightweight filter strategy and fine-tuning strategy that refines the anchor with negligible overhead. AnchorDS produces finer-grained detail, more natural colours, and stronger semantic consistency, particularly for complex prompts, while maintaining efficiency. Extensive experiments show that our method surpasses previous methods in both quality and efficiency.

文本到3D生成稳定性动态锚点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。