用大模型反馈提升3D生成的文本一致性,解决多物交互难题。
CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback
- 引入多模态大模型反馈优化文本-3D对齐,改进传统SDS方法
- 在TIFA数据集上多个指标显著提升,实现更一致的3D生成结果
- 适合关注3D生成文本对齐与多物体交互的研究者
Score Distillation Sampling(SDS)在文本到3D内容生成中取得显著进展,但基于SDS的方法在涉及多个物体及其复杂交互时,难以保持用户提示的语义保真度。现有方法通常通过在3D数据集上微调多视角扩散模型来增强3D一致性,但这反而加剧了文本-3D对齐的退化。问题根源在于SDS在优化过程中累积了视角无关的偏差,逐渐偏离理想文本对齐方向。为此,我们提出一种新目标——文本一致得分蒸馏(TCSD),利用多模态大语言模型(MLLM)的跨模态理解能力,在优化过程中评估并引导文本-3D对应关系。我们进一步开发了3DLLaVA-CRITIC,一个专用于评估3D生成中多视角文本对齐的微调MLLM。此外,我们引入基于LLM的布局初始化,通过语义感知的空间配置显著加速优化收敛。所提出的CoherenDream框架在TIFA子集上多个指标均实现一致提升。作为首个将MLLM融入SDS优化的研究,我们还进行了大量消融实验,探索适用于3D生成任务的最优MLLM适配策略。
原文摘要 · Abstract (English)
Score Distillation Sampling (SDS) has achieved remarkable success in text-to-3D content generation. However, SDS-based methods struggle to maintain semantic fidelity for user prompts, particularly when involving multiple objects with intricate interactions. While existing approaches often address 3D consistency through multiview diffusion model fine-tuning on 3D datasets, this strategy inadvertently exacerbates text-3D alignment degradation. The limitation stems from SDS's inherent accumulation of view-independent biases during optimization, which progressively diverges from the ideal text alignment direction. To alleviate this limitation, we propose a novel SDS objective, dubbed as Textual Coherent Score Distillation (TCSD), which integrates alignment feedback from multimodal large language models (MLLMs). Our TCSD leverages cross-modal understanding capabilities of MLLMs to assess and guide the text-3D correspondence during the optimization. We further develop 3DLLaVA-CRITIC - a fine-tuned MLLM specialized for evaluating multiview text alignment in 3D generations. Additionally, we introduce an LLM-layout initialization that significantly accelerates optimization convergence through semantic-aware spatial configuration. Our framework, CoherenDream, achieves consistent improvement across multiple metrics on TIFA subset.As the first study to incorporate MLLMs into SDS optimization, we also conduct extensive ablation studies to explore optimal MLLM adaptations for 3D generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。