arXiv:2506.22056cs.AI2025-06中稿 · ICML被引 5

构建多模态轨迹检索框架,提升智能体行为建模能力

Universal Retrieval for Multimodal Trajectory Modeling

  • 用视觉语言模型结合对比学习优化轨迹检索
  • 在多个数据集上召回率优于现有基线方法
  • 适合研究智能体行为建模与多模态检索的学者

轨迹数据涵盖人类行为与环境状态的多模态信息,对提升人工智能代理能力具有重要意义,尤其在图形用户界面环境中。然而,随着轨迹数据的爆炸式增长,如何系统建模轨迹级表示仍面临挑战。本文提出多模态轨迹检索,连接通用检索与代理中心的轨迹建模。我们基于真实世界场景中的标注演示与状态构建了统一代理轨迹数据集(UATD),并推出包含大量轨迹检索对的基准测试GAE-Bench。此外,提出GAE-Retriever多模态检索框架,采用视觉语言模型并引入令牌选择与GradCache机制优化对比学习。在多个数据集上的全面评估表明,GAE-Retriever在检索召回率上持续超越强基线,验证了其在推进多模态轨迹检索方面的有效性。

原文摘要 · Abstract (English)

Trajectory data, capturing human actions and environmental states across various modalities, holds significant potential for enhancing AI agent capabilities, particularly in GUI environments. However, how to model the representation of trajectory-level data presents a significant challenge that has not been systematically addressed amid explosive trajectory data growth. In this work, we introduce Multimodal Trajectory Retrieval, bridging the gap between universal retrieval and agent-centric trajectory modeling. We construct the Unified Agent Trajectory Dataset (UATD) from annotated demonstrations and states across diverse real-world scenarios. Based on this, we present GAE-Bench, a benchmark containing a large number of trajectory-based retrieval pairs. In addition, we propose GAE-Retriever, a multimodal retrieval framework that adopts vision-language models and incorporates optimized contrastive learning through a token selection and the GradCache mechanism. Comprehensive evaluations across multiple datasets show that GAE-Retriever consistently outperforms strong baselines in retrieval recall, highlighting its effectiveness in advancing multimodal trajectory retrieval.

轨迹建模多模态检索智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。