arXiv:2508.05064cs.GRcs.CL2025-08被引 1

将语言嵌入与3D高斯点云结合,实现文本驱动的场景理解与生成。

A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding

  • 用语言嵌入引导高斯点云生成,实现文本控制的3D场景重建。
  • 整合大语言模型提升语义理解,支持交互式编辑和跨模态生成。
  • 适合研究多模态3D生成、智能机器人场景理解的开发者参考。

高斯点云(Gaussian Splatting)已成为实时3D场景表示的革新性技术,相比神经辐射场(NeRF)在效率与表达力上更具优势,已推动场景重建、机器人导航及交互内容创作等领域发展。近期,将大语言模型(LLMs)与语言嵌入融入高斯点云流程,为文本条件生成、编辑及语义场景理解开辟新路径。然而,该交叉领域的系统性综述仍显不足。本文对当前融合语言引导与3D高斯点云的研究进行结构化梳理,涵盖理论基础、集成策略及真实应用场景,指出计算瓶颈、泛化能力弱以及缺乏语义标注的3D高斯数据等关键局限,并提出未来发展方向,以推进语言引导的3D场景理解与生成技术进展。

原文摘要 · Abstract (English)

Gaussian Splatting has rapidly emerged as a transformative technique for real-time 3D scene representation, offering a highly efficient and expressive alternative to Neural Radiance Fields (NeRF). Its ability to render complex scenes with high fidelity has enabled progress across domains such as scene reconstruction, robotics, and interactive content creation. More recently, the integration of Large Language Models (LLMs) and language embeddings into Gaussian Splatting pipelines has opened new possibilities for text-conditioned generation, editing, and semantic scene understanding. Despite these advances, a comprehensive overview of this emerging intersection has been lacking. This survey presents a structured review of current research efforts that combine language guidance with 3D Gaussian Splatting, detailing theoretical foundations, integration strategies, and real-world use cases. We highlight key limitations such as computational bottlenecks, generalizability, and the scarcity of semantically annotated 3D Gaussian data and outline open challenges and future directions for advancing language-guided 3D scene understanding using Gaussian Splatting.

3D生成语言嵌入高斯点云多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。