arXiv:2607.24298cs.CV2026-07

让3D生成模型在多图输入下稳定表现,通过动态选择最相关图像。

UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing

论文配图:UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focus Cross-Attention Routing
图 1 · 摘自论文原文
  • 每个3D体素在去噪时自动选择最匹配的输入图
  • 在无约束多图输入下显著减少几何扭曲和纹理过平
  • 无需训练,可直接接入现有单图3D模型

当前3D基础模型虽能从单图生成高质量资产,但在无约束多图输入下性能急剧下降,常出现几何失真、纹理过平和颜色混乱。我们认为问题不在于模型容量不足,而在于单图交叉注意力机制与多图场景不匹配:现有模型缺乏在去噪每一步决定每个3D体素应信任哪张图的合理机制。重新审视单图3D模型后,我们发现显式引导每个体素至最具信息量的图像即可显著提升不一致多图输入下的表现。基于此,提出Umi3D——一种无需训练、即插即用的多图3D生成框架。其核心Simultaneous Focus Cross-Attention(SFC-Attn)在每步去噪中激活所有条件图像,同时允许每个体素聚焦于最能解释它的单张图像。为此,我们推导出体素-图像亲和性度量Voxel Reference Score(VRS),无需外部匹配、分割或对应模型。大量实验表明,Umi3D可充分释放单图3D生成框架在多样化任务中的多图潜力。

原文摘要 · Abstract (English)

Recent 3D foundation models can generate high-quality assets from a single image, but degrade markedly on unconstrained multi-image inputs, often producing distorted geometry, over-smoothed textures, and chaotic colors. We argue that this failure stems not from limited model capacity, but from a mismatch between single-image cross-attention and the multi-image setting: existing models lack a principled way to decide which image each 3D voxel should trust at each denoising step. Revisiting recent single-image 3D foundation models, we show that explicitly routing each voxel to its most informative image is sufficient to unlock strong performance on inconsistent multi-image inputs. Based on this observation, we propose UMI3D, a training-free and plug-and-play framework that restructures cross-attention for unconstrained multi-image 3D generation. Its core, Simultaneous Focus Cross-Attention (SFC-Attn), activates all conditioning images at each denoising step while allowing each voxel to focus on the single image that best explains it. To enable this routing, we derive the Voxel Reference Score (VRS), a model-intrinsic metric for voxel--image affinity that requires no external matching, segmentation, or correspondence models. Extensive experiments show that UMI3D unlocks the multi-image potential of single-image 3D generation frameworks across diverse tasks. Project Page: UMI3D-Project.github.io.

3D生成多图输入注意力路由扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。