构建三元关系结构,让机器人在新场景中更懂如何操作物体。
TriRelVLA: Triadic Relational Structure for Generalizable Embodied Manipulation

- 用物体-手-任务三元关系显式建模交互
- 跨场景、跨物体、跨任务泛化性能显著提升
- 适合需要强泛化能力的机器人操作研究
视觉-语言-动作(VLA)模型在训练过的任务上表现良好,但在未见场景和物体上泛化能力差。根本原因在于其隐式视觉表征将物体外观、背景和场景布局纠缠在一起,导致策略对视觉变化敏感。现有方法虽通过结构化中间表示提升可迁移性,但多关注场景语义而非与动作相关的关联。本文观察到操作行为取决于物体-手-任务的三元关系结构,该结构决定任务需求、机器人状态与物体属性之间的交互。为此提出TriRelVLA框架:1)从多模态输入构建显式的物体-手-任务三元关系表征作为关系基元;2)构建任务引导的关系图,使用任务感知交叉注意力生成节点,关系感知图变换器建模节点间交互;3)进行关系条件下的动作生成,将关系结构压缩至瓶颈空间并投影至大语言模型以预测动作。该三元关系瓶颈降低对外观统计的依赖,实现跨场景、跨物体及跨任务组合的泛化。此外引入真实机器人数据集用于微调。实验表明,在微调任务上表现优异,并在跨场景、跨物体、跨任务泛化上均有明显提升。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models perform well on training-seen robotic tasks but struggle to generalize to unseen scenes and objects. A key limitation lies in their implicit visual representations, which entangle object appearance, background, and scene layout. This makes policies sensitive to visual variations. Prior work improves transferability through structured intermediate representations that objectify visual content. However, these representations mainly capture scene semantics instead of action-relevant relations. As a result, action prediction remains tied to appearance statistics. We observe that manipulation actions depend on the object-hand-task relational structure, which governs interactions among task requirements, robot states, and object properties. Based on this observation, we propose TriRelVLA, a triadic relational VLA framework for generalizable embodied manipulation. Our approach consists of three components: 1) We construct explicit object-hand-task triadic representations from multimodal inputs as relational primitives. 2) We build a task-grounded relational graph. Task-guided cross-attention forms nodes, and a relation-aware graph transformer models interactions among them. 3) We perform relation-conditioned action generation. The relational structure is compressed into a bottleneck space and projected into the LLM for action prediction. This triadic relational bottleneck reduces reliance on appearance statistics and enables transfer across scenes, objects, and task compositions. We further introduce a real-world robotic dataset for fine-tuning. Experiments show strong performance on fine-tuned tasks and clear gains in cross-scene, cross-object, and cross-task generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。