用向量量化特征场实现快速语义提升,支持文本驱动的局部场景编辑。
Vector Quantized Feature Fields for Fast 3D Semantic Lifting
- 引入向量量化特征场,按需高效检索像素对齐的语义掩码。
- 在复杂室内外场景中实现高效语义提升,显著加速智能体问答效率。
- 适合需要轻量级语义理解与交互的场景建模应用。
我们通过引入每视角的掩码来扩展传统提升方法至语义提升,这些掩码指示了用于提升任务的相关像素。掩码通过查询多尺度像素对齐的特征图获得,该特征图源自如压缩特征场和特征点云等场景表示。然而,存储从压缩特征场渲染的每视角特征图不切实际,而特征点云则存储和查询成本高昂。为实现轻量级按需检索像素对齐的相关性掩码,我们提出向量量化特征场(Vector-Quantized Feature Field)。我们在复杂室内与室外场景中验证了其有效性。结合向量量化特征场的语义提升可解锁多种场景表示与具身智能应用。具体而言,我们展示了该方法如何实现文本驱动的局部场景编辑,并显著提升具身问答的效率。
原文摘要 · Abstract (English)
We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding multiscale pixel-aligned feature maps, which are derived from scene representations such as distilled feature fields and feature point clouds. However, storing per-view feature maps rendered from distilled feature fields is impractical, and feature point clouds are expensive to store and query. To enable lightweight on-demand retrieval of pixel-aligned relevance masks, we introduce the Vector-Quantized Feature Field. We demonstrate the effectiveness of the Vector-Quantized Feature Field on complex indoor and outdoor scenes. Semantic lifting, when paired with a Vector-Quantized Feature Field, can unlock a myriad of applications in scene representation and embodied intelligence. Specifically, we showcase how our method enables text-driven localized scene editing and significantly improves the efficiency of embodied question answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。