解决工业级文生3D生成的领域适配与几何推理难题
ForgeDreamer: Industrial Text-to-3D Generation with Multi-Expert LoRA and Cross-View Hypergraph
- 用多专家LoRA集成实现跨类别知识融合,避免干扰
- 通过跨视图超图建模,捕捉多视角结构依赖关系
- 适合需要高精度制造的工业3D生成场景
当前文生3D方法在自然场景表现良好,但在工业应用中面临两大瓶颈:传统LoRA融合导致跨类别知识干扰,且成对一致性约束无法捕捉精密制造所需的高阶结构依赖。我们提出ForgeDreamer框架,通过两项创新解决上述问题。首先,引入多专家LoRA集成机制,将多个类别专属LoRA模型统一为表示,实现优异的跨类别泛化能力,同时消除知识干扰。其次,基于增强语义理解,构建跨视图超图几何增强方法,同步捕获多视角间的结构依赖关系。二者协同提升语义理解与几何推理能力,超图建模保障制造级一致性。在自建工业数据集上的大量实验表明,该方法在语义泛化和几何保真度上均优于现有最先进方法。代码已开源。
原文摘要 · Abstract (English)
Current text-to-3D generation methods excel in natural scenes but struggle with industrial applications due to two critical limitations: domain adaptation challenges where conventional LoRA fusion causes knowledge interference across categories, and geometric reasoning deficiencies where pairwise consistency constraints fail to capture higher-order structural dependencies essential for precision manufacturing. We propose a novel framework named ForgeDreamer addressing both challenges through two key innovations. First, we introduce a Multi-Expert LoRA Ensemble mechanism that consolidates multiple category-specific LoRA models into a unified representation, achieving superior cross-category generalization while eliminating knowledge interference. Second, building on enhanced semantic understanding, we develop a Cross-View Hypergraph Geometric Enhancement approach that captures structural dependencies spanning multiple viewpoints simultaneously. These components work synergistically improved semantic understanding, enables more effective geometric reasoning, while hypergraph modeling ensures manufacturing-level consistency. Extensive experiments on a custom industrial dataset demonstrate superior semantic generalization and enhanced geometric fidelity compared to state-of-the-art approaches. Code is available at https://github.com/Junhaocai27/ForgeDreamer
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。