arXiv:2604.14914cs.CV2026-04中稿 · ECCV被引 1

破解3D生成模型对文本指令失灵的难题,实现不受限形状的精准语义编辑。

Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

论文配图:Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
图 1 · 摘自论文原文
  • 通过解耦几何表达与语言敏感性,绕过文本引导失效的隐空间陷阱。
  • 在未见文本条件下仍能生成复杂几何形状,保持高保真度。
  • 适合需要跨域3D形状语义操作的研究者与创作者。

基于文本的生成模型逆向是操控2D或3D内容的核心范式,广泛应用于文本编辑、风格迁移及逆问题求解。然而,该方法依赖生成模型对自然语言提示保持敏感的假设。我们发现,对于最先进的原生文本到3D生成模型,这一假设常失效。我们识别出关键失败模式:生成轨迹被吸引至隐空间“陷落区”,即模型对提示修改不再敏感的区域,在这些区域中,输入文本变化无法有效改变内部表征以影响输出几何结构。关键在于,这并非模型几何表达能力的局限;相同模型具备生成丰富多样形状的能力,但对分布外文本指导变得不敏感。我们通过分析生成模型的采样轨迹揭示此现象,并发现可通过利用模型的无条件生成先验,仍可表示并生成复杂几何。由此提出更鲁棒的文本驱动3D编辑框架,通过解耦几何表达力与语言敏感性,避开隐空间陷阱。该方法克服现有3D流水线局限,实现对分布外3D形状的高保真语义操纵。

原文摘要 · Abstract (English)

Text-driven inversion of generative models is a core paradigm for manipulating 2D or 3D content, unlocking numerous applications such as text-based editing, style transfer, or inverse problems. However, it relies on the assumption that generative models remain sensitive to natural language prompts. We demonstrate that for state-of-the-art native text-to-3D generative models, this assumption often collapses. We identify a critical failure mode where generation trajectories are drawn into latent "sink traps": regions where the model becomes insensitive to prompt modifications. In these regimes, changes to the input text fail to alter internal representations in a way that alters the output geometry. Crucially, we observe that this is not a limitation of the model's \textit{geometric} expressivity; the same generative models possess the ability to produce a vast diversity of shapes but, as we demonstrate, become insensitive to out-of-distribution \textit{text} guidance. We investigate this behavior by analyzing the sampling trajectories of the generative model, and find that complex geometries can still be represented and produced by leveraging the model's unconditional generative prior. This leads to a more robust framework for text-based 3D shape editing that bypasses latent sinks by decoupling a model's geometric representation power from its linguistic sensitivity. Our approach addresses the limitations of current 3D pipelines and enables high-fidelity semantic manipulation of out-of-distribution 3D shapes. Project webpage: https://daidedou.sorpi.fr/publication/beyondprompts

3D生成文本控制逆向生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。