arXiv:2412.09008cs.CVcs.HC2024-12被引 5

用手绘+语音在元宇宙中快速生成3D模型,20秒完成。

MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments

  • 手绘草图结合语音指令,用ControlNet生成逼真图像。
  • 生成高质量3D网格耗时不足20秒,支持实时交互。
  • 适合元宇宙创作、设计原型等需要快速建模的场景。

我们提出MS2Mesh-XR,一种新颖的多模态手绘到3D网格生成流程,支持用户在扩展现实(XR)环境中通过空中手绘草图并结合语音输入创建逼真的3D物体。用户可在虚拟空间中自然地进行空中手绘,系统通过集成语音输入,利用ControlNet根据草图和解析后的文本提示生成真实感图像。用户可审阅并选择偏好图像,随后使用卷积重建模型将其转换为高细节3D网格。该流程可在20秒内生成高质量3D网格,支持运行时XR场景中的沉浸式可视化与操作。我们通过两个XR应用场景验证了该方法的实用性。借助自然用户输入与前沿生成式AI能力,本方法显著提升了基于XR的创意生产效率,优化了用户体验。代码与演示地址:https://yueqiu0911.github.io/MS2Mesh-XR/

原文摘要 · Abstract (English)

We present MS2Mesh-XR, a novel multi-modal sketch-to-mesh generation pipeline that enables users to create realistic 3D objects in extended reality (XR) environments using hand-drawn sketches assisted by voice inputs. In specific, users can intuitively sketch objects using natural hand movements in mid-air within a virtual environment. By integrating voice inputs, we devise ControlNet to infer realistic images based on the drawn sketches and interpreted text prompts. Users can then review and select their preferred image, which is subsequently reconstructed into a detailed 3D mesh using the Convolutional Reconstruction Model. In particular, our proposed pipeline can generate a high-quality 3D mesh in less than 20 seconds, allowing for immersive visualization and manipulation in run-time XR scenes. We demonstrate the practicability of our pipeline through two use cases in XR settings. By leveraging natural user inputs and cutting-edge generative AI capabilities, our approach can significantly facilitate XR-based creative production and enhance user experiences. Our code and demo will be available at: https://yueqiu0911.github.io/MS2Mesh-XR/

3D生成手绘建模XR交互多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。