让多个3D物体自然互动且多视角一致生成。
Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation

- 用交互碰撞机制引导高斯点云合理分布,实现物理可信的物体交互。
- 多视角自适应分数蒸馏采样提升跨视图一致性,避免视觉幻觉。
- 支持灵活编辑,适合复杂场景的高质量3D资产生成。
近期3D生成技术借助文本到图像扩散模型取得显著进展,但现有方法仍面临两大挑战:(1) 主要生成单个3D物体,难以构建多物体组合场景,因缺乏对高斯原始体在合理交互区域的建模;(2) 在3D优化过程中常出现跨视图不一致问题,因分数蒸馏采样(Score Distillation Sampling)在单视图上独立执行,不可避免引发跨视图幻觉。为此,我们提出I2C-3D,一种基于优化的新方法,可生成多视角一致、具有合理交互的组合式3D资产。具体而言,我们设计了包含式交互碰撞策略,引导高斯原始体自然出现在合理交互区域,确保组合场景中物体间具备物理合理性与视觉连贯性。此外,为增强多视角一致性,提出多视角自适应分数蒸馏采样,通过调节实例令牌与空间令牌的注意力图,从预训练扩散模型中蒸馏多视图一致性先验与布局先验。得益于上述设计,I2C-3D不仅生成高保真、多视角一致的组合3D资产,还支持灵活3D编辑,助力复杂场景生成。大量实验表明,I2C-3D在生成质量与多视图一致性上均优于现有方法。
原文摘要 · Abstract (English)
Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) They primarily generate single 3D object, but struggle to generate multi-object compositional 3D assets due to the lack of the modeling for Gaussian primitives in reasonable interactions. (2) They often suffer from cross-view inconsistency during 3D optimization, as Score Distillation Sampling inherently performs on each single view, inevitably resulting in cross-view hallucinations. To solve above issues, we propose I2C-3D, a novel optimization-based method to generate multi-view consistent compositional 3D assets with reasonable interactions. Specifically, we propose an Inclusive Interactive Collisions strategy to guide Gaussian primitives appearing in reasonable interaction regions naturally, thereby ensuring objects in the compositional scene interact in a physically plausible and visually coherent way. Additionally, to enhance multi-view consistency, Multi-View Adaptive Score Distillation Sampling is devised to distill multi-view consistency prior and layout prior from pre-trained diffusion model by modulating attention map of instance token and spatial token across viewpoints. Benefiting from above elaborate designs, I2C-3D not only generates high-fidelity multi-view consistent compositional 3D assets but also supports 3D editing flexibly, facilitating complex scene generation. Extensive experiments demonstrate our I2C-3D outperforms existing methods in generation quality and multi-view consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。