让机器人理解任务结构,提升复杂操作成功率
Differentiate-and-Inject: Enhancing VLAs via Functional Differentiation Induced by In-Parameter Structural Reasoning
- 将任务结构嵌入模型参数,实现无需提示的分步推理
- 在多个操作基准上成功率显著优于现有方法
- 适合需要灵活任务分解的机器人控制场景
随着机器人需执行更复杂的任务,其不仅需理解低层动作,还需掌握决定任务展开方式的高层结构。现有视觉-语言-动作(VLA)模型在任务级推理方面存在局限:或依赖不稳定的提示式上下文分解,或通过端到端长时程训练,但后者需大规模示范且混淆了任务推理与底层控制。本文提出在参数中嵌入任务结构的iSTAR框架,通过参数空间内的结构推理实现功能分化。该方法将任务语义结构以隐式动态场景图形式注入模型参数,捕捉物体关系、子任务语义和任务依赖。在多个操作基准测试中,iSTAR展现出更可靠的任务分解能力与更高的成功率达78.3%,验证了参数空间结构推理在功能分化与泛化上的有效性。
原文摘要 · Abstract (English)
As robots are expected to perform increasingly diverse tasks, they must understand not only low-level actions but also the higher-level structure that determines how a task should unfold. Existing vision-language-action (VLA) models struggle with this form of task-level reasoning. They either depend on prompt-based in-context decomposition, which is unstable and sensitive to linguistic variations, or end-to-end long-horizon training, which requires large-scale demonstrations and entangles task-level reasoning with low-level control. We present in-parameter structured task reasoning (iSTAR), a framework for enhancing VLA models via functional differentiation induced by in-parameter structural reasoning. Instead of treating VLAs as monolithic policies, iSTAR embeds task-level semantic structure directly into model parameters, enabling differentiated task-level inference without external planners or handcrafted prompt inputs. This injected structure takes the form of implicit dynamic scene-graph knowledge that captures object relations, subtask semantics, and task-level dependencies in parameter space. Across diverse manipulation benchmarks, iSTAR achieves more reliable task decompositions and higher success rates than both in-context and end-to-end VLA baselines, demonstrating the effectiveness of parameter-space structural reasoning for functional differentiation and improved generalization across task variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。