一个模型搞定不同机器人追踪,零样本适配新平台。
AdaTracker: Learning Adaptive In-Context Policy for Cross-Embodiment Active Visual Tracking

- 用上下文编码器动态捕捉机器人物理特性,自适应调整控制策略。
- 在仿真和真实场景中均实现跨平台零样本追踪,性能显著优于现有方法。
- 适合需要快速部署到多种机器人的视觉追踪任务,尤其看重泛化能力的场景。
在不同机器人平台上实现统一的主动视觉追踪极具挑战性,因各平台物理约束与运动动力学差异显著。现有方法通常为每种机器人训练独立模型,导致可扩展性差、泛化能力有限。为此,我们提出AdaTracker,一种基于上下文自适应的策略学习框架,可在多样化机器人形态上稳健追踪目标。核心思想是通过实体上下文编码器显式建模实体特异性约束,从历史数据中推断出该约束,并动态调制上下文感知策略,使其能在零样本条件下为未见实体生成最优控制动作。为提升鲁棒性,引入两个辅助目标以确保上下文识别准确性和时序一致性。在仿真与真实世界中的实验表明,AdaTracker在跨实体泛化、样本效率及零样本适应方面显著优于当前最先进方法。
原文摘要 · Abstract (English)
Realizing active visual tracking with a single unified model across diverse robots is challenging, as the physical constraints and motion dynamics vary drastically from one platform to another. Existing approaches typically train separate models for each embodiment, leading to poor scalability and limited generalization. To address this, we propose AdaTracker, an adaptive in-context policy learning framework that robustly tracks targets on diverse robot morphologies. Our key insight is to explicitly model embodiment-specific constraints through an Embodiment Context Encoder, which infers embodiment-specific constraints from history. This contextual representation dynamically modulates a Context-Aware Policy, enabling it to infer optimal control actions for unseen embodiments in a zero-shot manner. To enhance robustness, we introduce two auxiliary objectives to ensure accurate context identification and temporal consistency. Experiments in both simulation and the real world demonstrate that AdaTracker significantly outperforms state-of-the-art methods in cross-embodiment generalization, sample efficiency, and zero-shot adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。