让智能体状态信息更精准地融入大模型,提升任务理解与决策能力。
CLSP: High-Fidelity Contrastive Language-State Pre-training for Agent State Representation
- 用对比学习融合语言与状态信息,构建高保真表征
- 在文本-状态检索等任务上超越基线,提升导航与理解性能
- 适合强化学习与多模态大模型研究者参考
随着人工智能快速发展,多模态学习成为重要方向。对智能体而言,状态是传递精确信息的关键模态,与图像、视频、语言等共同构成感知基础。尤其在强化学习和多模态大语言模型广泛应用的背景下,状态表征仍存在不足。为此,本文提出高保真对比语言-状态预训练(CLSP)方法,可将状态信息准确编码为通用表示,适用于强化学习与多模态大语言模型。首先设计基于分类的预训练任务,用粗粒度信息训练编码器;随后构建状态与语言描述的数据对,利用预训练编码器初始化CLSP编码器;再通过对比学习训练,有效表达精确状态信息。此外,采用随机傅里叶特征(RFF)增强数值信息表示,实现高保真映射。大量实验表明,该表示具备优异精度与泛化能力,在文本-状态检索、强化学习导航任务及多模态大模型理解中表现突出。
原文摘要 · Abstract (English)
With the rapid development of artificial intelligence, multimodal learning has become an important research area. For intelligent agents, the state is a crucial modality to convey precise information alongside common modalities like images, videos, and language. This becomes especially clear with the broad adoption of reinforcement learning and multimodal large language models. Nevertheless, the representation of state modality still lags in development. To this end, we propose a High-Fidelity Contrastive Language-State Pre-training (CLSP) method, which can accurately encode state information into general representations for both reinforcement learning and multimodal large language models. Specifically, we first design a pre-training task based on the classification to train an encoder with coarse-grained information. Next, we construct data pairs of states and language descriptions, utilizing the pre-trained encoder to initialize the CLSP encoder. Then, we deploy contrastive learning to train the CLSP encoder to effectively represent precise state information. Additionally, we enhance the representation of numerical information using the Random Fourier Features (RFF) method for high-fidelity mapping. Extensive experiments demonstrate the superior precision and generalization capabilities of our representation, achieving outstanding results in text-state retrieval, reinforcement learning navigation tasks, and multimodal large language model understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。