通过姿态引导注意力,提升机器人动作生成的精准与稳定
PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- 用姿态条件约束视觉注意力,聚焦任务相关区域
- 在多个机器人操作基准上实现更精准、高效的动作序列
- 轻量设计无需额外感知模块,适合实时场景应用
视觉-语言-动作(VLA)模型在具身任务中表现优异,但在复杂环境中仍易生成冗余或不稳定的动作,影响其在时间敏感场景中的应用。本文认为,现有VLA的空间均匀感知场会分散注意力至无关目标。为此,提出PosA-VLA框架,通过姿态条件化锚定注意力机制,持续引导模型聚焦任务相关视觉区域。该机制使指令语义与可行动视觉线索更好对齐,显著提升动作生成的精度与效率。框架采用轻量结构,无需分割或定位等辅助模块,保障高效推理。大量实验表明,该方法在多个机器人操纵基准上均表现出高精度、低延迟的行为,并在多种挑战性环境中展现强泛化能力。
原文摘要 · Abstract (English)
The Vision-Language-Action (VLA) models have demonstrated remarkable performance on embodied tasks and shown promising potential for real-world applications. However, current VLAs still struggle to produce consistent and precise target-oriented actions, as they often generate redundant or unstable motions along trajectories, limiting their applicability in time-sensitive scenarios.In this work, we attribute these redundant actions to the spatially uniform perception field of existing VLAs, which causes them to be distracted by target-irrelevant objects, especially in complex environments.To address this issue, we propose an efficient PosA-VLA framework that anchors visual attention via pose-conditioned supervision, consistently guiding the model's perception toward task-relevant regions. The pose-conditioned anchor attention mechanism enables the model to better align instruction semantics with actionable visual cues, thereby improving action generation precision and efficiency. Moreover, our framework adopts a lightweight architecture and requires no auxiliary perception modules (e.g., segmentation or grounding networks), ensuring efficient inference. Extensive experiments verify that our method executes embodied tasks with precise and time-efficient behavior across diverse robotic manipulation benchmarks and shows robust generalization in a variety of challenging environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。