arXiv:2603.23202cs.CV2026-03被引 3

用人类注视模式指导机器人视觉注意力,提升操作精度与可解释性。

Gaze-Regularized Vision-Language-Action Models for Robotic Manipulation

  • 将人眼注视热图转为像素级分布,通过KL散度约束模型注意力
  • 在多个任务上提升4-12%性能,减少训练步数且抗光照噪声
  • 无需眼动设备,适用于现有数据集,适合需高可信度的机器人场景

尽管视觉-语言-动作(VLA)模型取得进展,机器人执行精细操作仍受限于缺乏主动视觉注意分配机制。人类注视自然反映了意图、规划与执行模式,是引导机器人感知的有力监督信号。本文提出一种注视正则化训练框架,无需架构修改或推理开销,即可使VLA模型内部注意力对齐人类视觉模式。方法将时序聚合的注视热图转化为块级分布,通过KL散度正则化Transformer注意力,引入对任务相关特征的归纳偏置,同时保持部署效率。集成至现有VLA架构后,在多个操作基准测试中实现4-12%的性能提升。注视正则化模型以更少训练步骤达到相当性能,并在光照变化与传感器噪声下保持鲁棒性。此外,学习到的注意力模式生成可解释可视化,反映人类策略,增强对机器人的信任。本框架无需眼动设备,可直接应用于现有数据集。结果表明,人类感知先验能显著加速机器人学习,同时提升任务表现与系统可解释性。

原文摘要 · Abstract (English)

Despite advances in Vision-Language-Action (VLA) models, robotic manipulation struggles with fine-grained tasks because current models lack mechanisms for active visual attention allocation. Human gaze naturally encodes intent, planning, and execution patterns -- offering a powerful supervisory signal for guiding robot perception. We introduce a gaze-regularized training framework that aligns VLA models' internal attention with human visual patterns without architectural modifications or inference-time overhead. Our method transforms temporally aggregated gaze heatmaps into patch-level distributions and regularizes the transformer's attention through KL divergence, creating an inductive bias toward task-relevant features while preserving deployment efficiency. When integrated into existing VLA architectures, our approach yields 4-12% improvements across manipulation benchmarks. The gaze-regularized models reach equivalent performance with fewer training steps and maintain robustness under lighting variations and sensor noise. Beyond performance metrics, the learned attention patterns produce interpretable visualizations that mirror human strategies, enhancing trust in robotic systems. Moreover, our framework requires no eye-tracking equipment and applies directly to existing datasets. These results demonstrate that human perceptual priors can significantly accelerate robot learning while improving both task performance and system interpretability.

机器人操作注意力机制可解释性多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。