用多模态查询精准匹配3D资产,让虚拟场景更连贯
MetaFind: Scene-Aware 3D Asset Retrieval for Coherent Metaverse Scene Generation
- 支持文本、图像、3D混合查询,联合建模物体与场景结构
- 在多个任务中显著提升空间与风格一致性,优于基线方法
- 适合构建连贯元宇宙场景的开发者与设计师使用
我们提出MetaFind,一种面向元宇宙场景生成的场景感知三模态组合检索框架,用于从大规模3D资产库中检索高质量3D资产。针对现有方法在空间、语义和风格约束下检索不一致,以及缺乏专门面向3D资产检索的标准范式两大挑战,MetaFind创新性地设计了灵活的检索机制,支持任意组合的文本、图像和3D模态作为查询。通过联合建模对象级特征(包括外观)与场景级布局结构,增强空间推理与风格一致性。方法上引入可插拔的等变布局编码器ESSGNN,捕捉空间关系与物体外观特征,确保检索资产在坐标变换下仍与场景上下文及风格一致。框架支持迭代式场景构建,持续根据当前场景更新调整检索结果。实证评估表明,相比基线方法,MetaFind在多种检索任务中显著提升了空间与风格一致性。
原文摘要 · Abstract (English)
We present MetaFind, a scene-aware tri-modal compositional retrieval framework designed to enhance scene generation in the metaverse by retrieving 3D assets from large-scale repositories. MetaFind addresses two core challenges: (i) inconsistent asset retrieval that overlooks spatial, semantic, and stylistic constraints, and (ii) the absence of a standardized retrieval paradigm specifically tailored for 3D asset retrieval, as existing approaches mainly rely on general-purpose 3D shape representation models. Our key innovation is a flexible retrieval mechanism that supports arbitrary combinations of text, image, and 3D modalities as queries, enhancing spatial reasoning and style consistency by jointly modeling object-level features (including appearance) and scene-level layout structures. Methodologically, MetaFind introduces a plug-and-play equivariant layout encoder ESSGNN that captures spatial relationships and object appearance features, ensuring retrieved 3D assets are contextually and stylistically coherent with the existing scene, regardless of coordinate frame transformations. The framework supports iterative scene construction by continuously adapting retrieval results to current scene updates. Empirical evaluations demonstrate the improved spatial and stylistic consistency of MetaFind in various retrieval tasks compared to baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。