arXiv:2411.14322cs.ROcs.CV2024-11被引 4

用3D高斯点云实现视觉重排任务,让智能体直接比对图像状态。

SplatR : Experience Goal Visual Rearrangement with 3D Gaussian Splatting and Dense Feature Matching

  • 用3D高斯点云表示场景,支持快速渲染真实视图。
  • 在AI2-THOR上优于现有方法,提升重排成功率。
  • 适合做具身智能中的场景重建与状态对比任务。

体验目标视觉重排任务是具身人工智能中的基础挑战,要求智能体构建一个准确的环境模型以恢复被打乱的场景至原始状态。本文提出一种新框架,利用3D高斯点云作为3D场景表示,实现高质量、高保真的新视角快速渲染。该方法使智能体能够同时获得当前状态与目标状态的一致视图,从而在图像空间中直接比较两者的差异。为实现精准匹配,我们采用基于基础模型提取的密集特征匹配方法,利用其通用特征表示能力,增强鲁棒性与泛化性。我们在AI2-THOR重排基准上验证了该方法,结果表明其性能超越现有最优方法。

原文摘要 · Abstract (English)

Experience Goal Visual Rearrangement task stands as a foundational challenge within Embodied AI, requiring an agent to construct a robust world model that accurately captures the goal state. The agent uses this world model to restore a shuffled scene to its original configuration, making an accurate representation of the world essential for successfully completing the task. In this work, we present a novel framework that leverages on 3D Gaussian Splatting as a 3D scene representation for experience goal visual rearrangement task. Recent advances in volumetric scene representation like 3D Gaussian Splatting, offer fast rendering of high quality and photo-realistic novel views. Our approach enables the agent to have consistent views of the current and the goal setting of the rearrangement task, which enables the agent to directly compare the goal state and the shuffled state of the world in image space. To compare these views, we propose to use a dense feature matching method with visual features extracted from a foundation model, leveraging its advantages of a more universal feature representation, which facilitates robustness, and generalization. We validate our approach on the AI2-THOR rearrangement challenge benchmark and demonstrate improvements over the current state of the art methods

具身智能3D重建视觉重排高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。