解决3D生成中因视角偏见导致的多视图不一致问题
Debiasing Diffusion Priors via 3D Attention for Consistent Gaussian Splatting

- 引入3D感知注意力引导与分层调制模块,增强跨视角一致性
- 在多个3D任务中显著提升多视图一致性,效果优于基线模型
- 适合需要精准3D编辑与生成的研究者使用
基于文本到图像(T2I)扩散模型的通用3D任务(如生成或编辑)因无需大量3D训练数据而受到广泛关注。然而,T2I模型存在先验视角偏见,导致物体不同视角间外观冲突。该偏见使主题词在跨注意力(CA)计算中优先激活先验视角特征,无视目标视角条件。本文通过数学分析揭示其根源,并发现UNet不同层对视角偏见的影响不同。为此提出TD-Attn框架,包含两个核心组件:(1) 3D-Aware Attention Guidance Module(3D-AAG)构建视图一致的3D注意力高斯,强化注意力区域的空间一致性,弥补单视图CA图空间信息不足;(2) Hierarchical Attention Modulation Module(HAM)利用语义引导树(SGT)指导语义响应分析器(SRP),定位并调制对视角敏感的CA层,优化后的CA图进一步支持更一致的3D注意力高斯构建。实验表明,TD-Attn可作为通用插件,显著提升多种3D任务的多视图一致性。
原文摘要 · Abstract (English)
Versatile 3D tasks (e.g., generation or editing) that distill from Text-to-Image (T2I) diffusion models have attracted significant research interest for not relying on extensive 3D training data. However, T2I models exhibit limitations resulting from prior view bias, which produces conflicting appearances between different views of an object. This bias causes subject-words to preferentially activate prior view features during cross-attention (CA) computation, regardless of the target view condition. To overcome this limitation, we conduct a comprehensive mathematical analysis to reveal the root cause of the prior view bias in T2I models. Moreover, we find different UNet layers show different effects of prior view in CA. Therefore, we propose a novel framework, TD-Attn, which addresses multi-view inconsistency via two key components: (1) the 3D-Aware Attention Guidance Module (3D-AAG) constructs a view-consistent 3D attention Gaussian for subject-words to enforce spatial consistency across attention-focused regions, thereby compensating for the limited spatial information in 2D individual view CA maps; (2) the Hierarchical Attention Modulation Module (HAM) utilizes a Semantic Guidance Tree (SGT) to direct the Semantic Response Profiler (SRP) in localizing and modulating CA layers that are highly responsive to view conditions, where the enhanced CA maps further support the construction of more consistent 3D attention Gaussians. Notably, HAM facilitates semantic-specific interventions, enabling controllable and precise 3D editing. Extensive experiments firmly establish that TD-Attn has the potential to serve as a universal plugin, significantly enhancing multi-view consistency across 3D tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。