语音直接生成3D模型,让AR设备实时高效运行。
From Voices to Worlds: Developing an AI-Powered Framework for 3D Object Generation in Augmented Reality
- 用语音指令驱动3D生成,结合大语言模型理解上下文。
- 通过简化网格降低文件大小,提升资源受限设备的响应速度。
- 适合教育、设计、无障碍应用,开源促进跨领域创新。
本文提出Matrix框架,用于增强现实(AR)环境中的实时3D物体生成。该系统融合先进的文本到3D生成模型、多语言语音转文字翻译及大语言模型(LLMs),支持用户通过语音指令进行交互。系统可处理语音输入,生成3D对象并基于上下文提供推荐,显著提升AR体验。核心优势在于通过减少网格复杂度优化3D模型,大幅降低文件大小与计算开销,从而在资源受限的AR设备上实现更快处理。框架还配备预生成物体库,进一步减轻GPU负载。本研究展示了其在教育、设计与无障碍领域的应用潜力,并展望未来支持图像转3D、环境物体检测和多模态输入。框架开源,推动跨行业持续创新。
原文摘要 · Abstract (English)
This paper presents Matrix, an advanced AI-powered framework designed for real-time 3D object generation in Augmented Reality (AR) environments. By integrating a cutting-edge text-to-3D generative AI model, multilingual speech-to-text translation, and large language models (LLMs), the system enables seamless user interactions through spoken commands. The framework processes speech inputs, generates 3D objects, and provides object recommendations based on contextual understanding, enhancing AR experiences. A key feature of this framework is its ability to optimize 3D models by reducing mesh complexity, resulting in significantly smaller file sizes and faster processing on resource-constrained AR devices. Our approach addresses the challenges of high GPU usage, large model output sizes, and real-time system responsiveness, ensuring a smoother user experience. Moreover, the system is equipped with a pre-generated object repository, further reducing GPU load and improving efficiency. We demonstrate the practical applications of this framework in various fields such as education, design, and accessibility, and discuss future enhancements including image-to-3D conversion, environmental object detection, and multimodal support. The open-source nature of the framework promotes ongoing innovation and its utility across diverse industries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。