arXiv:2606.05975cs.CVcs.RO2026-06被引 1

让机器人根据任务需求,快速定位3D场景中物体的功能部件。

T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation

论文配图:T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation
图 1 · 摘自论文原文
  • 基于任务描述筛选场景图中的相关物体,分层定位功能部件。
  • 在SceneFun3D上达到顶尖性能,速度更快、内存占用更低。
  • 适合需要高效感知的机器人应用,如智能家务与交互。

开放词汇3D功能分割使机器人能够在3D场景中定位物体的功能组件,该任务需兼具空间理解与任务解析能力。现有方法多聚焦于物体级识别,或对场景进行全范围零件分割,导致资源消耗大、耗时长。为缓解这一矛盾,我们提出T-FunS3D,一种任务驱动的层次化开放词汇3D功能分割方法,旨在为机器人应用提供可操作的感知能力。输入包括室内场景的3D点云和带姿态的RGB-D图像。通过提取环境中的实例及其视觉嵌入,构建开放词汇场景图。给定任务描述后,T-FunS3D在场景图中识别最相关实例,并利用视觉语言模型定位其功能部件。在SceneFun3D数据集上的实验表明,T-FunS3D在开放词汇3D功能分割性能上媲美当前最优方法,同时实现更快速的运行时间和更低的内存占用。

原文摘要 · Abstract (English)

Open-vocabulary 3D functionality segmentation enables robots to localize functional object components in 3D scenes. It is a challenging task that requires spatial understanding and task interpretation. Current open-vocabulary 3D segmentation methods primarily focus on object-level recognition, while scene-wide part segmentation methods attempt to segment the entire scene exhaustively, making them highly resource-intensive and time consuming. Balancing segmentation performance in terms of granularity, accuracy, and speed remains a challenge. As one step towards alleviating this, we introduce T-FunS3D, a task-driven hierarchical open-vocabulary 3D functionality segmentation method that provides actionable perception for robotic applications. Our method takes as input the 3D point cloud and posed RGB-D images of an indoor scene. We construct an open-vocabulary scene graph by extracting instances and their visual embeddings in the environment. Given a task description, T-FunS3D identifies the most relevant instances in the scene graph and locates their functional components leveraging a vision-language model. Experiments on the SceneFun3D dataset demonstrate that T-FunS3D is comparable to state-of-the-art in open-vocabulary 3D functionality segmentation, while achieving faster runtime and reduced memory usage.

3D分割机器人感知视觉语言模型任务驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。