统一框架实现3D检索与可控4D生成,提升跨模态对齐精度。
A Three-Level Alignment Framework for Large-Scale 3D Retrieval and Controlled 4D Generation
- 通过三级对齐机制融合文本、3D模型与图像,优化语义匹配。
- 在Align3D 130数据集上实现高精度3D检索与时间一致的4D生成。
- 适合需要可控3D/4D内容生成的研究者与工业应用开发者。
我们提出Uni4D,一个基于文本、3D模型与图像三模态结构化三级对齐的统一框架,支持大规模开放词汇3D检索与可控4D生成。基于Align3D 130数据集,Uni4D采用3D文本多头注意力与搜索模型,通过增强语义对齐优化文本到3D的检索性能。框架进一步通过三个组件强化跨模态对齐:精确的文本到3D检索、多视角3D到图像对齐,以及用于生成时序一致4D资产的图像到文本对齐。实验表明,Uni4D在3D检索与可控4D生成任务中均取得高质量结果,推动动态多模态理解与实际应用发展。
原文摘要 · Abstract (English)
We introduce Uni4D, a unified framework for large scale open vocabulary 3D retrieval and controlled 4D generation based on structured three level alignment across text, 3D models, and image modalities. Built upon the Align3D 130 dataset, Uni4D employs a 3D text multi head attention and search model to optimize text to 3D retrieval through improved semantic alignment. The framework further strengthens cross modal alignment through three components: precise text to 3D retrieval, multi view 3D to image alignment, and image to text alignment for generating temporally consistent 4D assets. Experimental results demonstrate that Uni4D achieves high quality 3D retrieval and controllable 4D generation, advancing dynamic multimodal understanding and practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。