arXiv:2604.08410cs.CVcs.RO2026-04

让机器人听懂指令并精准执行复杂抓取操作

BLaDA: Bridging Language to Functional Dexterous Actions within 3DGS Fields

论文配图:BLaDA: Bridging Language to Functional Dexterous Actions within 3DGS Fields
图 1 · 摘自论文原文
  • 用语言解析生成结构化操作约束,实现语义与动作的精准匹配
  • 通过三角几何约束定位功能区域,提升3D空间推理一致性
  • 零样本适配多种任务,适合需要高精度灵巧操作的研究场景

在非结构化环境中,功能性灵巧抓取需要语义理解、精确的3D功能定位和物理可解释的执行。模块化分层方法比端到端视觉-语言-动作(VLA)方法更可控且可解释,但现有方法仍依赖预定义的可用性标签,缺乏语义与姿态间的紧密耦合。为此,我们提出BLaDA(在3DGS场中连接语言与灵巧动作),一个可解释的零样本框架,将开放词汇指令转化为感知与控制约束,用于功能性灵巧操作。BLaDA通过知识引导的语言解析(KLP)模块将自然语言解析为结构化的六元组操作约束,建立可解释的推理链。为实现姿态一致的空间推理,引入三角功能点定位(TriLocation)模块,利用3D高斯泼溅作为连续场景表示,并在三角几何约束下识别功能区域。最后,3D关键点抓取矩阵变换执行(KGT3D+)模块将这些语义-几何约束解码为物理上合理的腕部姿态和手指级指令。在复杂基准上的大量实验表明,BLaDA在可用性定位精度和跨多种类别与任务的功能操作成功率方面显著优于现有方法。代码将在https://github.com/PopeyePxx/BLaDA公开。

原文摘要 · Abstract (English)

In unstructured environments, functional dexterous grasping calls for the tight integration of semantic understanding, precise 3D functional localization, and physically interpretable execution. Modular hierarchical methods are more controllable and interpretable than end-to-end VLA approaches, but existing ones still rely on predefined affordance labels and lack the tight semantic--pose coupling needed for functional dexterous manipulation. To address this, we propose BLaDA (Bridging Language to Dexterous Actions in 3DGS fields), an interpretable zero-shot framework that grounds open-vocabulary instructions as perceptual and control constraints for functional dexterous manipulation. BLaDA establishes an interpretable reasoning chain by first parsing natural language into a structured sextuple of manipulation constraints via a Knowledge-guided Language Parsing (KLP) module. To achieve pose-consistent spatial reasoning, we introduce the Triangular Functional Point Localization (TriLocation) module, which utilizes 3D Gaussian Splatting as a continuous scene representation and identifies functional regions under triangular geometric constraints. Finally, the 3D Keypoint Grasp Matrix Transformation Execution (KGT3D+) module decodes these semantic-geometric constraints into physically plausible wrist poses and finger-level commands. Extensive experiments on complex benchmarks demonstrate that BLaDA significantly outperforms existing methods in both affordance grounding precision and the success rate of functional manipulation across diverse categories and tasks. Code will be publicly available at https://github.com/PopeyePxx/BLaDA.

灵巧操作3D高斯零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。