arXiv:2606.29828cs.CV2026-06AAAI被引 2

用多视角图像实现室内物体零样本定制生成,细节更真实。

HomeDiffusion: Zero-Shot Object Customization with Multi-View Representation Learning for Indoor Scenes

论文配图:HomeDiffusion: Zero-Shot Object Customization with Multi-View Representation Learning for Indoor Scenes
图 1 · 摘自论文原文
  • 基于多视角图像构建参考对象表征,提升生成一致性
  • 在扩散过程中跨注意力保留参考物细节,生成更逼真物体
  • 适合电商、家装等需快速定制物品视觉效果的场景

近期,零样本物体定制生成方法迅速发展,展现出巨大应用潜力。例如,在电商领域,用户可预览家具摆放在自己家中的视觉效果或衣物穿在身上的样子。现有方法多基于扩散模型和提取的参考物体特征进行定制生成,但生成物在图案、曲线等细节上与原参考物存在显著偏差。尤其对于非对称物体,因缺乏全面的多视角信息,难以生成与背景场景协调的物体姿态。为此,我们构建了一个包含家具与室内场景多角度图像的新数据集。基于扩散模型,提出HomeDiffusion,利用同一参考物体的多视角图像,在指定区域精准生成与背景和谐的物体姿态。在扩散过程中,进一步提取参考物体的高保真细节,并与潜在空间中的噪声隐变量进行交叉注意力,确保定制生成中细节的保留。大量定性和定量实验表明,该方法在零样本及少样本物体定制生成任务上均优于现有方法。

原文摘要 · Abstract (English)

Recently, zero-shot object customization generation methods have rapidly developed and shown tremendous potential for applications. For instance, in the e-commerce domain, consumers can observe the visual effect of furniture placed within their personal living spaces or clothes worn on their own bodies. Many existing approaches perform object customization generation based on diffusion models and extracted reference object features. However, the generated object significantly diverges from the original reference object in details such as patterns and curves. Particularly for asymmetrical reference objects, the absence of comprehensive multi-viewpoint information prevents the generation of object poses that harmonize with the background scene. To address these shortcomings, we have constructed a novel dataset comprising multi-angle images of furniture and indoor scenes. Based on diffusion models, we introduce HomeDiffusion, which can leverage multi-viewpoint images of the same reference object to accurately generate visually harmonious object poses within specified areas of the background scene. During the diffusion process, we further extract high-fidelity details of the reference object and perform cross-attention with the noise latents in the latent space, thereby ensuring the preservation of details in the customized object generation. Extensive qualitative and quantitative experiments demonstrate that our method achieves superior performance over other existing zero-shot as well as few-shot object customization approaches.

零样本生成扩散模型多视角表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。