arXiv:2601.07516cs.CLcs.AI2026-01ACL

用隐式动作空间提升多模态对话模型的强化学习效率

Controlling Multimodal Conversational Agents with Coverage-Enhanced Latent Actions

  • 构建紧凑隐式动作空间,结合图文与纯文本数据增强覆盖
  • 在两个对话任务上优于主流基线方法,适配多种强化学习算法
  • 适合研究多模态交互与强化学习融合的开发者

视觉-语言模型正被广泛用于多模态对话代理(MCAs)以应对多样化对话任务。近期,强化学习(RL)被广泛探索用于适应不同人机交互场景。尽管显著提升了泛化性能,通过RL微调MCAs仍面临处理极大文本词元空间的挑战。为此,我们学习一个紧凑的隐式动作空间用于RL微调。具体地,采用观察学习机制构建隐式动作空间的码本,利用未来观测估计当前可预测的隐式动作,进而用于重构未来观测。然而,成对图文数据稀缺限制了码本的充分覆盖。因此,我们结合成对图文数据与纯文本数据构建隐式动作空间,使用跨模态投影器将文本嵌入转换为图文嵌入。先在成对图文数据上初始化投影器,再在大量纯文本数据上通过新颖的循环一致性损失进一步训练,以增强鲁棒性。实验表明,该隐式动作方法在两个对话任务上均优于竞争性基线,在多种RL算法下表现更优。

原文摘要 · Abstract (English)

Vision-language models are increasingly employed as multimodal conversational agents (MCAs) for diverse conversational tasks. Recently, reinforcement learning (RL) has been widely explored for adapting MCAs to various human-AI interaction scenarios. Despite showing great enhancement in generalization performance, fine-tuning MCAs via RL still faces challenges in handling the extremely large text token space. To address this, we learn a compact latent action space for RL fine-tuning instead. Specifically, we adopt the learning from observation mechanism to construct the codebook for the latent action space, where future observations are leveraged to estimate current latent actions that could further be used to reconstruct future observations. However, the scarcity of paired image-text data hinders learning a codebook with sufficient coverage. Thus, we leverage both paired image-text data and text-only data to construct the latent action space, using a cross-modal projector for transforming text embeddings into image-text embeddings. We initialize the cross-modal projector on paired image-text data, and further train it on massive text-only data with a novel cycle consistency loss to enhance its robustness. We show that our latent action based method outperforms competitive baselines on two conversation tasks across various RL algorithms.

多模态对话强化学习隐式动作视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。