arXiv:2507.23772cs.CV2025-07被引 3

让机器人理解复杂场景中连续操作指令的3D功能区域。

SeqAffordSplat: Scene-level Sequential Affordance Reasoning on 3D Gaussian Splatting

  • 用语言模型生成带分割标记的文本,指导3D掩码逐步生成。
  • 在1800多个场景上实现序列化功能推理,性能超越现有方法。
  • 融合2D视觉模型语义特征,提升复杂场景理解能力。

3D功能推理任务旨在将人类指令与3D物体的功能区域关联,是具身智能体的关键能力。当前基于3D高斯点云(3DGS)的方法仅支持单对象、单步交互,难以应对真实世界中的长时序、多对象任务。为此,本文提出序列化3D高斯功能推理新任务,并构建包含1800+场景的大规模基准数据集SeqAffordSplat,以支持复杂3DGS环境下的长时序功能理解研究。我们提出端到端框架SeqSplatNet,直接将指令映射为3D功能掩码序列。该模型利用大语言模型自回归生成包含特殊分割标记的文本,引导条件解码器生成对应3D掩码。为处理复杂场景几何结构,引入条件几何重建预训练策略,使模型从已知几何观测中学习完整功能区域掩码重建,建立稳健几何先验。此外,设计特征注入机制,将2D视觉基础模型(VFM)的丰富语义特征分层融合至3D解码器,缓解语义歧义。大量实验表明,本方法在挑战性基准上达到新最优性能,成功将功能推理从单步交互推进至场景级复杂序列任务。

原文摘要 · Abstract (English)

3D affordance reasoning, the task of associating human instructions with the functional regions of 3D objects, is a critical capability for embodied agents. Current methods based on 3D Gaussian Splatting (3DGS) are fundamentally limited to single-object, single-step interactions, a paradigm that falls short of addressing the long-horizon, multi-object tasks required for complex real-world applications. To bridge this gap, we introduce the novel task of Sequential 3D Gaussian Affordance Reasoning and establish SeqAffordSplat, a large-scale benchmark featuring 1800+ scenes to support research on long-horizon affordance understanding in complex 3DGS environments. We then propose SeqSplatNet, an end-to-end framework that directly maps an instruction to a sequence of 3D affordance masks. SeqSplatNet employs a large language model that autoregressively generates text interleaved with special segmentation tokens, guiding a conditional decoder to produce the corresponding 3D mask. To handle complex scene geometry, we introduce a pre-training strategy, Conditional Geometric Reconstruction, where the model learns to reconstruct complete affordance region masks from known geometric observations, thereby building a robust geometric prior. Furthermore, to resolve semantic ambiguities, we design a feature injection mechanism that lifts rich semantic features from 2D Vision Foundation Models (VFM) and fuses them into the 3D decoder at multiple scales. Extensive experiments demonstrate that our method sets a new state-of-the-art on our challenging benchmark, effectively advancing affordance reasoning from single-step interactions to complex, sequential tasks at the scene level.

3D推理序列理解具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。