arXiv:2512.04513cs.AI2025-12

让多模态大模型与世界模型双向互动,提升智能体跨任务泛化能力。

BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models

  • 通过正反双向路径连接语言模型与环境模型的语义和动态空间
  • 在多任务和跨环境测试中表现优于现有方法,稳定性更强
  • 适合研究开放世界智能体、多模态决策的学者与开发者

构建通用具身智能体需要统一系统来理解多模态目标、建模环境动态,并在多样真实任务中执行可靠动作。多模态大语言模型(MLLMs)具备强语义先验和跨模态泛化能力,而世界模型(WMs)则提供可操作的潜在动态以支持预测与控制。二者结合有望实现开放式具身智能,但面临两大挑战:(1) 如何建立MLLM的语义意图与WM潜在空间中动态状态之间的紧密耦合;(2) 实现任务感知的自适应能力,支持多任务学习与跨环境泛化。为此,我们提出BiTAgent,一个任务感知的动态联合框架,实现MLLM与WM间的双向耦合。该框架包含两条互补路径:前向路径将MLLM表示注入WM潜在空间,实现语义引导的想象;后向路径则通过密集文本条件奖励,使WM生成的反馈优化MLLM的语义空间。这一双向交互由三个协同组件实现:任务感知动态联合学习、任务感知行为学习、以及MLLM-WM联合优化,共同协调语义推理与动态预测。在多任务与跨环境设置下的大量实验表明,其性能显著优于当前最优基线,标志着迈向开放式具身学习的重要一步。

原文摘要 · Abstract (English)

Building generalist embodied agents requires a unified system that can interpret multimodal goals, model environment dynamics, and execute reliable actions across diverse real-world tasks. Multimodal large language models (MLLMs) offer strong semantic priors and cross-modal generalization, while world models (WMs) provide actionable latent dynamics for prediction and control. Their combination holds promise for open-ended embodied intelligence, yet introduces two key challenges: (1) establishing a tight coupling between the semantic intent from MLLMs and the dynamic state representations within the WM's latent space, and (2) achieving task-aware adaptability that supports multi-task learning and cross-environment generalization. To address these limitations, we propose BiTAgent, a task-aware dynamic joint framework that enables bidirectional coupling between MLLMs and WMs. BiTAgent establishes two complementary pathways: a forward path that injects MLLM representations into the WM's latent space for semantically guided imagination, and a backward path where WM-generated feedback refines the MLLM's semantic space via dense text-conditioned rewards. This bidirectional interaction is realized through three synergistic components: Task-Aware Dynamic Joint Learning, Task-Aware Behavior Learning, and MLLM-WM Joint Optimization, which together harmonize semantic reasoning and dynamic prediction. Extensive experiments across multi-task and cross-environment settings demonstrate superior stability and generalization over state-of-the-art baselines, marking a step toward open-ended embodied learning.

具身智能多模态世界模型双向耦合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。