arXiv:2606.27872cs.ROcs.AI2026-06中稿 · IJCAI被引 1

用动态注意力提升机器人长任务执行能力

S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation

论文配图:S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
图 1 · 摘自论文原文
  • 通过状态空间跟踪任务进度,动态调整视觉、语言、动作信息融合权重
  • 20亿参数模型在长任务上超越70亿参数模型,达当前最优表现
  • 适合需要持续适应复杂阶段的机器人操作研究者

视觉-语言-动作(VLA)模型在机器人操作中表现强劲,但在长时序任务中因累积误差而性能下降。这主要源于静态特征融合机制依赖固定权重组合视觉、语言和动作表征,无法随任务阶段变化自适应调整。为此,我们提出S²-VLA框架,引入状态空间引导的自适应注意力(SSGAA)机制。该机制维护一个信念状态以追踪任务进展,并生成动态门控权重,自适应融合三类互补信息:视觉特征用于空间感知,任务意图用于高层规划,时间动作序列用于执行一致性。这种动态融合使模型能随任务演进调整关注重点,契合不同阶段需求。尽管仅含20亿参数,S²-VLA在长时序操作基准(包括LIBERO与SimplerEnv)上持续优于更大规模的70亿参数模型,达到当前最优性能,凸显自适应特征融合对长时序机器人操作的重要性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their performance degrades significantly in long-horizon tasks due to cumulative error propagation. This limitation largely arises from static feature fusion mechanisms that rely on fixed weights to combine visual, language, and action representations, preventing the model from adapting to different phases of task execution. To address this limitation, we propose S$^2$-VLA, a framework that introduces a State-Space Guided Adaptive Attention (SSGAA) mechanism. SSGAA maintains a belief state that tracks task progression and generates dynamic gating weights to adaptively fuse information from three complementary sources visual features for spatial perception, task intents for high-level task planning, and temporal action sequences for execution consistency. This adaptive fusion allows the model to shift its focus throughout task execution, aligning with the evolving requirements of different task stages. Despite its compact 2B parameter size, S$^2$-VLA consistently outperforms larger 7B-scale models and achieves state-of-the-art performance on long-horizon manipulation benchmarks, including LIBERO and SimplerEnv. highlighting the importance of adaptive feature fusion for long-horizon robotic manipulation.

机器人操作自适应融合长时序任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。