让3D视频流支持语言交互,实现自由视角下的智能编辑与高帧率渲染。
DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming

- 动态双透明度设计,分离颜色与语言特征优化,提升渲染效率。
- 通过插值变形场降低时间冗余,支持低帧率视频升频至高帧率。
- 仅43KB平均帧大小,实现开放词汇语义分割与高质量重建。
3D高斯点阵(3DGS)已成为从多视角视频重构可流式传输的自由视角视频(FVV)的有前景范式。然而,基于3DGS的FVV通常缺乏用户交互与编辑能力,削弱了沉浸体验。近期研究通过知识蒸馏将CLIP的语言特征融入3DGS,实现了开放词汇查询并支持多种下游应用。但FVV对低帧尺寸与高帧率的严格要求,使现有语言高斯表示难以适用。本文提出DLGStream,一种新型语言嵌入式FVV表示方法,通过流式传输随时间变化的语言特征与高斯属性,支持4D环境交互、场景编辑与空间智能。具体地,提出双透明度动态语言高斯表示,为颜色与语言特征分别维护两个透明度属性,以缓解颜色与特征联合优化导致的性能下降。此外,引入基于插值的变形场,减少时间冗余,该场亦可用于4D帧插值,将低帧率FVV序列提升至高帧率。实验表明,DLGStream在开放词汇分割与重建质量上均表现优异,平均帧尺寸仅为43 KB。代码已公开于https://github.com/kkkzh/DLGStream。
原文摘要 · Abstract (English)
3D Gaussian Splatting~(3DGS) has emerged as a promising paradigm for reconstructing streamable free-viewpoint video~(FVV) from multi-view videos. However, 3DGS-based FVVs typically lack user interaction and editing capabilities, which diminishes the immersive experience. Recent research has integrated language features from CLIP into 3DGS via distillation, enabling open-vocabulary queries and supporting many downstream applications. Nevertheless, the stringent requirements of FVV, low frame size and high FPS, make current language Gaussian representations unsuitable for language-embedded FVV. In this paper, we propose DLGStream, a novel language-embedded FVV representation that streams time-varying language features alongside Gaussian attributes to support 4D environment interaction, scene editing, and spatial intelligence. Specifically, we propose a dual-opacity dynamic language Gaussian representation, which maintains two opacity attributes for color and language features to deal with performance degradation that occurs when colors and features are jointly optimized. Furthermore, we introduce an interpolation-based deformation field to reduce temporal redundancy. This deformation field can also be used for 4D frame interpolation, boosting FVV sequences from low to high FPS. Experimental results demonstrate that DLGStream achieves superior performance in both on open-vocabulary segmentation and reconstruction quality with an average frame size of merely 43 KB. The code is available on \href{https://github.com/kkkzh/DLGStream}{https://github.com/kkkzh/DLGStream}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。