arXiv:2603.23478cs.CV2026-03被引 1

让大模型主动看视频,精准定位3D场景中的可操作部件。

UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Functionality Segmentation

  • 用大模型当主动观察者,一次推理完成语义时空联合分析。
  • 在SceneFun3D上实现59.9%的相对mIoU提升,无需任务训练。
  • 适合需要零样本理解复杂交互指令的机器人与AR应用。

3D场景中的功能分割要求智能体将隐式的自然语言指令转化为细粒度交互元素的精确掩码。现有方法依赖碎片化流程,在任务解析初期存在视觉盲区。我们发现这些方法受限于单尺度、被动且基于启发式的关键帧选择。本文提出UniFunc3D,一种统一的无训练框架,将多模态大语言模型视为主动观察者。通过在单次前向传播中整合语义、时间与空间推理,UniFunc3D实现对任务分解的直接视觉证据支持。该方法引入粗到精的主动时空定位策略,使模型能自适应选择正确视频帧,并聚焦高细节交互区域,同时保留全局上下文以消除歧义。在SceneFun3D数据集上,UniFunc3D达到领先性能,相比无训练与有训练方法均有显著提升,相对mIoU提高59.9%,且无需任何特定任务训练。代码将在项目页面发布:https://jiaying.link/unifunc3d。

原文摘要 · Abstract (English)

Functionality segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements. Existing methods rely on fragmented pipelines that suffer from visual blindness during initial task parsing. We observe that these methods are limited by single-scale, passive and heuristic frame selection. We present UniFunc3D, a unified and training-free framework that treats the multimodal large language model as an active observer. By consolidating semantic, temporal, and spatial reasoning into a single forward pass, UniFunc3D performs joint reasoning to ground task decomposition in direct visual evidence. Our approach introduces active spatial-temporal grounding with a coarse-to-fine strategy. This allows the model to select correct video frames adaptively and focus on high-detail interactive parts while preserving the global context necessary for disambiguation. On SceneFun3D, UniFunc3D achieves state-of-the-art performance, surpassing both training-free and training-based methods by a large margin with a relative 59.9\% mIoU improvement, without any task-specific training. Code will be released on our project page: https://jiaying.link/unifunc3d.

3D分割多模态大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。