arXiv:2603.22899cs.RO2026-03

用少样本演示实现边缘设备上的快速工件姿态校正

Agile-VLA: Few-Shot Industrial Pose Rectification via Implicit Affordance Anchoring

  • 通过隐式可用性锚定,直接将视觉特征映射为控制动作
  • 10赫兹感知与50赫兹控制解耦,5次示例即完成复杂工件校正
  • 适合工业场景中资源受限的实时机器人操控应用

在资源受限的边缘平台部署视觉-语言-动作(VLA)模型时,高延迟语义推理与动态操作所需的高频控制之间存在根本矛盾。为此,本文提出Agile-VLA,一种针对边缘设备(如NVIDIA Jetson Orin Nano)上工业姿态重定向任务的分层框架。核心创新是隐式可用性锚定机制,直接将几何视觉线索(如质心和边缘关键点锚点)映射为结构化的参数化动作基元,显著降低闭环控制中对高延迟语义推理的依赖。通过异步双流架构解耦感知(10 Hz)与控制(50 Hz),有效缓解了边缘机器人学习中的频率不匹配问题。在标准6-自由度机械臂上的实验表明,Agile-VLA仅需5次示例,即可通过外在灵巧性实现复杂不规则工件的鲁棒姿态校正。

原文摘要 · Abstract (English)

Deploying Vision-Language-Action (VLA) models on resource-constrained edge platforms encounters a fundamental conflict between high-latency semantic inference and the high-frequency control required for dynamic manipulation. To address the challenge, this paper presents Agile-VLA, a hierarchical framework designed for industrial pose reorientation tasks on edge devices such as the NVIDIA Jetson Orin Nano. The core innovation is an Implicit Affordance Anchoring mechanism that directly maps geometric visual cues, specifically centroid and rim keypoint anchors, into structured parametric action primitives, thereby substantially reducing reliance on high-latency semantic inference during closed-loop control. By decoupling perception (10 Hz) from control (50 Hz) via an asynchronous dual-stream architecture, the system effectively mitigates the frequency mismatch inherent in edge-based robot learning. Experimental results on a standard 6-DoF manipulator demonstrate that Agile-VLA achieves robust rectification of complex, irregular workpieces using only 5-shot demonstrations through extrinsic dexterity.

机器人操控边缘计算少样本学习姿态校正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。