用AI把图片语音变3D模型,实时放进增强现实
Generative AI Framework for 3D Object Generation in Augmented Reality
- 输入图片或语音,通过VLM和LLM生成3D模型
- 支持多语言、复杂环境,实现毫秒级模型加载
- 适合游戏、零售、设计等需快速建模的场景
本论文提出一个集成先进生成式AI模型的框架,用于在增强现实(AR)环境中实时生成三维(3D)物体。核心目标是将图像、语音等多样化输入转化为精确的3D模型,提升用户交互与沉浸感。关键组件包括先进的目标检测算法、友好的交互技术,以及Shap-E等强大生成模型。系统利用视觉语言模型(VLMs)和大语言模型(LLMs),从图像中捕捉空间细节,处理文本信息,生成完整3D对象,并无缝融入真实环境。该框架在游戏、教育、零售和室内设计等领域有广泛应用:玩家可创建个性化游戏资产,顾客可在购买前预览产品在实际环境中的效果,设计师能将实物快速转为3D模型进行实时可视化。其重要贡献在于降低3D建模门槛,使先进AI工具惠及更广泛用户,激发创造力与创新。框架还解决了多语言输入、多样视觉数据及复杂环境下的挑战,提升了目标检测与模型生成的准确性,实现实时加载3D模型至AR空间。结论表明,该框架有效融合生成式AI与AR技术,推动3D建模高效化,增强用户交互体验。
原文摘要 · Abstract (English)
This thesis presents a framework that integrates state-of-the-art generative AI models for real-time creation of three-dimensional (3D) objects in augmented reality (AR) environments. The primary goal is to convert diverse inputs, such as images and speech, into accurate 3D models, enhancing user interaction and immersion. Key components include advanced object detection algorithms, user-friendly interaction techniques, and robust AI models like Shap-E for 3D generation. Leveraging Vision Language Models (VLMs) and Large Language Models (LLMs), the system captures spatial details from images and processes textual information to generate comprehensive 3D objects, seamlessly integrating virtual objects into real-world environments. The framework demonstrates applications across industries such as gaming, education, retail, and interior design. It allows players to create personalized in-game assets, customers to see products in their environments before purchase, and designers to convert real-world objects into 3D models for real-time visualization. A significant contribution is democratizing 3D model creation, making advanced AI tools accessible to a broader audience, fostering creativity and innovation. The framework addresses challenges like handling multilingual inputs, diverse visual data, and complex environments, improving object detection and model generation accuracy, as well as loading 3D models in AR space in real-time. In conclusion, this thesis integrates generative AI and AR for efficient 3D model generation, enhancing accessibility and paving the way for innovative applications and improved user interactions in AR environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。