用少量数据实现跨操作者稳定意图识别,提升远程操控效率
Adaptor: Advancing Assistive Teleoperation with Few-Shot Learning and Cross-Operator Generalization

- 通过噪声注入和关键帧提取处理轨迹,建模意图不确定性
- 融合视觉语言模型与专家网络,实现少样本高精度动作生成
- 在不同水平操作者间表现稳定,适合实际多用户场景
辅助式远程操控通过共享控制提升效率,但因操作者习惯与能力差异,导致轨迹分布高度异质,影响意图识别稳定性。本文提出Adaptor,一种少样本跨操作者意图识别框架。该方法分两阶段:(i) 预处理阶段,通过噪声注入合成轨迹扰动以建模意图不确定性,并进行几何感知的关键帧提取;(ii) 策略学习阶段,使用意图专家编码处理后的轨迹,并融合预训练视觉-语言模型上下文,条件化动作专家生成动作。在真实世界与仿真基准上的实验表明,Adaptor达到当前最优性能,显著提升成功率与效率;且在不同专业水平操作者间表现出低方差,验证了强跨操作者泛化能力。
原文摘要 · Abstract (English)
Assistive teleoperation enhances efficiency via shared control, yet inter-operator variability, stemming from diverse habits and expertise, induces highly heterogeneous trajectory distributions that undermine intent recognition stability. We present Adaptor, a few-shot framework for robust cross-operator intent recognition. The Adaptor bridges the domain gap through two stages: (i) preprocessing, which models intent uncertainty by synthesizing trajectory perturbations via noise injection and performs geometry-aware keyframe extraction; and (ii) policy learning, which encodes the processed trajectories with an Intention Expert and fuses them with the pre-trained vision-language model context to condition an Action Expert for action generation. Experiments on real-world and simulated benchmarks demonstrate that Adaptor achieves state-of-the-art performance, improving success rates and efficiency over baselines. Moreover, the method exhibits low variance across operators with varying expertise, demonstrating robust cross-operator generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。