用自然语言精准修改3D模型,不破坏未编辑部分。
Vinedresser3D: Agentic Text-guided 3D Editing
- 通过多模态大模型理解指令并分解为结构与外观编辑指引。
- 在3D隐空间中进行无掩码编辑,保持内容一致性与几何完整。
- 适合需要高质量、无需手动标注的3D内容创作者使用。
文本引导的3D编辑旨在通过自然语言指令修改现有3D资产。当前方法难以同时理解复杂提示、自动定位编辑区域并保留未编辑内容。我们提出Vinedresser3D,一种直接在原生3D生成模型隐空间中运行的智能体框架。给定3D资产与编辑指令,该框架利用多模态大语言模型推断原始资产的丰富描述,识别编辑区域与类型(添加、修改、删除),并生成分层的结构与外观级文本指导。智能体选择信息量高的视角,调用图像编辑模型获取视觉引导。最后,基于反向流重建的隐空间修补流水线结合交错采样模块,在3D隐空间执行编辑,确保指令对齐的同时维持3D连贯性与未编辑区域。在多样化3D编辑任务上的实验表明,Vinedresser3D在自动指标和人工偏好评估中均优于基线方法,实现精确、连贯且无需掩码的3D编辑。
原文摘要 · Abstract (English)
Text-guided 3D editing aims to modify existing 3D assets using natural-language instructions. Current methods struggle to jointly understand complex prompts, automatically localize edits in 3D, and preserve unedited content. We introduce Vinedresser3D, an agentic framework for high-quality text-guided 3D editing that operates directly in the latent space of a native 3D generative model. Given a 3D asset and an editing prompt, Vinedresser3D uses a multimodal large language model to infer rich descriptions of the original asset, identify the edit region and edit type (addition, modification, deletion), and generate decomposed structural and appearance-level text guidance. The agent then selects an informative view and applies an image editing model to obtain visual guidance. Finally, an inversion-based rectified-flow inpainting pipeline with an interleaved sampling module performs editing in the 3D latent space, enforcing prompt alignment while maintaining 3D coherence and unedited regions. Experiments on diverse 3D edits demonstrate that Vinedresser3D outperforms prior baselines in both automatic metrics and human preference studies, while enabling precise, coherent, and mask-free 3D editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。