arXiv:2409.16019cs.RO2024-09被引 10

用具身大模型提升3D重建效率,自动规划视角并修正错误。

AIR-Embodied: An Efficient Active 3DGS-based Interaction and Reconstruction Framework with Embodied Large Language Model

  • 用多模态提示理解当前重建状态,动态规划观察视角和交互动作。
  • 在虚拟与真实场景中均实现更高质量、更高效的3D重建结果。
  • 适合需要智能自主重建的机器人、AR/VR等应用开发者使用。

近期3D重建与神经渲染技术虽提升了数字资产质量,但现有方法在应对不同物体形状、纹理及遮挡时泛化能力差。尽管下一最佳视角(NBV)规划与学习方法提供了解决方案,却受限于预设准则,难以像人类一样处理遮挡问题。为此,我们提出AIR-Embodied框架,将具身AI代理与大规模预训练多模态语言模型结合,以改进主动式3DGS重建。该框架采用三阶段流程:通过多模态提示理解当前重建状态,规划视点选择与交互动作,利用闭环推理确保执行准确。智能体根据计划与实际结果间的差异动态调整行为。在虚拟与真实环境中的实验表明,AIR-Embodied显著提升了重建效率与质量,为主动3D重建挑战提供了稳健解决方案。

原文摘要 · Abstract (English)

Recent advancements in 3D reconstruction and neural rendering have enhanced the creation of high-quality digital assets, yet existing methods struggle to generalize across varying object shapes, textures, and occlusions. While Next Best View (NBV) planning and Learning-based approaches offer solutions, they are often limited by predefined criteria and fail to manage occlusions with human-like common sense. To address these problems, we present AIR-Embodied, a novel framework that integrates embodied AI agents with large-scale pretrained multi-modal language models to improve active 3DGS reconstruction. AIR-Embodied utilizes a three-stage process: understanding the current reconstruction state via multi-modal prompts, planning tasks with viewpoint selection and interactive actions, and employing closed-loop reasoning to ensure accurate execution. The agent dynamically refines its actions based on discrepancies between the planned and actual outcomes. Experimental evaluations across virtual and real-world environments demonstrate that AIR-Embodied significantly enhances reconstruction efficiency and quality, providing a robust solution to challenges in active 3D reconstruction.

3D重建具身智能多模态主动感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。