arXiv:2607.09260cs.CV2026-07

实现开放词汇的实时3D场景探索,支持语音交互与高质量重建。

AnythingReality: Robust Online Gaussian Splatting SLAM for Open-Vocabulary VR Scene Exploration

论文配图:AnythingReality: Robust Online Gaussian Splatting SLAM for Open-Vocabulary VR Scene Exploration
图 1 · 摘自论文原文
  • 结合ORB-SLAM3与在线高斯点云重建,适应噪声数据。
  • 在自建数据集上提升14.5% PSNR,TUM-RGBD上提升11.7% PSNR。
  • 支持语音控制、沉浸式浏览,适合虚拟现实应用开发。

我们提出一种新型集成架构,实现鲁棒的在线3D高斯点云重建、实时虚拟现实(VR)探索及语音驱动的视觉-语言模型交互。不同于依赖干净深度或外部位姿的方法,本系统融合基于ORB-SLAM3的位姿估计与在线高斯重建,适用于嘈杂的真实世界数据。一个VR管道支持对增量重建场景的沉浸式探索;语义模块可转录语音指令、生成场景描述并记录兴趣点。相比现有先进在线高斯点云方法,在自建数据集上提升14.5% PSNR、8.6% SSIM、降低14.3% LPIPS;在TUM-RGBD数据集上提升11.7% PSNR、7.8% SSIM、降低21.6% LPIPS,且帧率相当或更优。系统实现88%的视觉-语言模型物体识别率。

原文摘要 · Abstract (English)

We present a novel integrated architecture for robust online 3D Gaussian splatting, real-time VR exploration, and speech-driven Vision-Language-Model interaction. Unlike methods assuming clean depth or external poses, our system combines ORB-SLAM3-based pose estimation with online Gaussian reconstruction for noisy real-world data. A VR pipeline enables immersive exploration of incremental reconstructions; a semantic module transcribes voice commands, generates scene descriptions, and records points of interest. Against state-of-the-art online Gaussian splatting methods, we improve image quality on our dataset (+14.5% PSNR, +8.6% SSIM, -14.3% LPIPS) and TUM-RGBD (+11.7% PSNR, +7.8% SSIM, -21.6% LPIPS), with comparable or superior frame rates via quality-speed configurations. We achieve an 88% VLM object-recognition rate.

3D重建虚拟现实语音交互高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。