arXiv:2603.21295cs.CV2026-03被引 1

图文联合引导3D生成,提升细节与视角一致性。

Text-Image Conditioned 3D Generation

  • 双分支结构分别处理图像和文本,轻量融合跨模态信息。
  • 图文联合生成在多指标上优于单一模态模型,显著减少视角偏差。
  • 适合需要高保真、多视角一致3D内容的设计师与开发者。

高质量3D资产对虚拟现实、工业设计和娱乐至关重要,推动了从用户提示生成3D内容的生成模型发展。现有3D生成器多依赖单一条件模态:图像引导模型利用像素对齐线索实现高视觉保真度,但输入视图有限或模糊时易产生视角偏差;文本引导模型提供广泛语义指导,却缺乏低层视觉细节。这限制了用户表达意图的能力。本研究发现,即使简单的晚期融合图文预测结果,也优于单模态模型,揭示出跨模态互补性。为此,我们提出文本-图像联合3D生成任务,并引入TIGON——一个双分支基线模型,包含独立的图像与文本编码器及轻量级跨模态融合机制。大量实验表明,图文联合条件可稳定超越单模态方法,凸显视觉-语言协同引导在3D生成中的潜力。

原文摘要 · Abstract (English)

High-quality 3D assets are essential for VR/AR, industrial design, and entertainment, motivating growing interest in generative models that create 3D content from user prompts. Most existing 3D generators, however, rely on a single conditioning modality: image-conditioned models achieve high visual fidelity by exploiting pixel-aligned cues but suffer from viewpoint bias when the input view is limited or ambiguous, while text-conditioned models provide broad semantic guidance yet lack low-level visual detail. This limits how users can express intent and raises a natural question: can these two modalities be combined for more flexible and faithful 3D generation? Our diagnostic study shows that even simple late fusion of text- and image-conditioned predictions outperforms single-modality models, revealing strong cross-modal complementarity. We therefore formalize Text-Image Conditioned 3D Generation, which requires joint reasoning over a visual exemplar and a textual specification. To address this task, we introduce TIGON, a minimalist dual-branch baseline with separate image- and text-conditioned backbones and lightweight cross-modal fusion. Extensive experiments show that text-image conditioning consistently improves over single-modality methods, highlighting complementary vision-language guidance as a promising direction for future 3D generation research. Project page: https://jumpat.github.io/tigon-page

3D生成图文联合跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。