arXiv:2603.22846cs.AI2026-03被引 4

用对抗竞争训练追踪智能体,提升复杂场景下的跟踪能力。

CoMaTrack: Competitive Multi-Agent Game-Theoretic Tracking with Vision-Language-Action Models

  • 构建多智能体对抗环境,通过竞争机制训练追踪与反追踪策略。
  • 30亿参数视觉语言模型在挑战性数据集上达92.1%精准率,超越70亿参数单智能体方法。
  • 首个基于Habitat的开放基准,支持动态对抗场景下的标准化评估。

具身视觉追踪(EVT)是具身智能中的核心动态任务,要求智能体根据语言指令精确追踪目标。现有方法多依赖单智能体模仿学习,受限于昂贵的专家数据和静态训练环境,泛化能力不足。受竞争驱动能力演化的启发,我们提出CoMaTrack,一种基于博弈论的多智能体强化学习框架,在动态对抗环境中训练智能体,实现更强的自适应规划与抗干扰策略。我们进一步推出CoMaTrack-Bench,首个基于Habitat的开源基准协议与剧集集,支持语言条件下的动态对战追踪,涵盖多样环境与指令,可实现主动对抗交互下的标准化鲁棒性评估。实验表明,CoMaTrack在标准基准与CoMaTrack-Bench上均达当前最优。值得注意的是,使用本框架训练的30亿参数视觉语言模型(VLM)在挑战性EVT-Bench上超越基于70亿参数模型的单智能体模仿学习方法,取得92.1%的STT、74.2%的DT和57.5%的AT。基准代码将开源于https://github.com/wlqcode/CoMaTrack-Bench。

原文摘要 · Abstract (English)

Embodied Visual Tracking (EVT), a core dynamic task in embodied intelligence, requires an agent to precisely follow a language-specified target. Yet most existing methods rely on single-agent imitation learning, suffering from costly expert data and limited generalization due to static training environments. Inspired by competition-driven capability evolution, we propose CoMaTrack, a competitive game-theoretic multi-agent reinforcement learning framework that trains agents in a dynamic adversarial setting with competitive subtasks, yielding stronger adaptive planning and interference-resilient strategies. We further introduce CoMaTrack-Bench, the first open-source Habitat-based benchmark protocol and episode set for language-conditioned competitive EVT featuring dynamic dueling, featuring game scenarios between a tracker and adaptive opponents across diverse environments and instructions, enabling standardized robustness evaluation under active adversarial interactions. Experiments show that CoMaTrack achieves state-of-the-art results on both standard benchmarks and CoMaTrack-Bench. Notably, a 3B VLM trained with our framework surpasses previous single-agent imitation learning methods based on 7B models on the challenging EVT-Bench, achieving 92.1% in STT, 74.2% in DT, and 57.5% in AT. The benchmark code will be available at https://github.com/wlqcode/CoMaTrack-Bench.

多智能体视觉追踪博弈论具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。