arXiv:2504.08581cs.CV2025-04被引 3

让3D场景中的物体部件能用自然语言精准查询并交互

FMLGS: Fast Multilevel Language Embedded Gaussians for Part-level Interactive Agents

  • 基于SAM2构建多层级语义映射,实现物体与部件的统一语义编码
  • 在部分级查询上准确率领先,速度比现有方法快2.5至98倍
  • 可嵌入对话系统,支持用户通过聊天与3D场景实时互动

语义交互辐射场是实现具身智能场景理解与操作的重要基础,但多粒度交互仍面临语言模糊性及部件级查询质量下降的挑战。本文提出FMLGS,支持在3D高斯溅射(3DGS)中进行部件级开放词汇查询。通过基于Segment Anything Model 2(SAM2)的高效管道,构建并查询一致的对象与部件级语义。设计语义偏移策略缓解部件间语言歧义,通过细粒度目标语义特征插值增强信息表达。训练后,可使用自然语言查询对象及其可描述部件。对比实验表明,本方法不仅能更精准定位部件级目标,且在速度与精度上均达领先水平:相比LERF快98倍,比LangSplat快4倍,比LEGaussians快2.5倍。进一步将FMLGS集成为虚拟代理,可通过聊天界面实现3D场景导航、目标定位与需求响应,展现出广阔的应用前景。

原文摘要 · Abstract (English)

The semantically interactive radiance field has long been a promising backbone for 3D real-world applications, such as embodied AI to achieve scene understanding and manipulation. However, multi-granularity interaction remains a challenging task due to the ambiguity of language and degraded quality when it comes to queries upon object components. In this work, we present FMLGS, an approach that supports part-level open-vocabulary query within 3D Gaussian Splatting (3DGS). We propose an efficient pipeline for building and querying consistent object- and part-level semantics based on Segment Anything Model 2 (SAM2). We designed a semantic deviation strategy to solve the problem of language ambiguity among object parts, which interpolates the semantic features of fine-grained targets for enriched information. Once trained, we can query both objects and their describable parts using natural language. Comparisons with other state-of-the-art methods prove that our method can not only better locate specified part-level targets, but also achieve first-place performance concerning both speed and accuracy, where FMLGS is 98 x faster than LERF, 4 x faster than LangSplat and 2.5 x faster than LEGaussians. Meanwhile, we further integrate FMLGS as a virtual agent that can interactively navigate through 3D scenes, locate targets, and respond to user demands through a chat interface, which demonstrates the potential of our work to be further expanded and applied in the future.

3D生成自然语言交互高斯溅射具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。