arXiv:2410.15885cs.AI2024-10EMNLP被引 6

让AI同时聊天和做决策,性能远超传统方法

VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making

论文配图:VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making
图 1 · 摘自论文原文
  • 采用多输入多输出架构,避免任务间互相干扰
  • 在自动驾驶平台测试中显著超越现有模型
  • 适合需要并行处理多任务的智能系统

近期的大规模预训练模型(如GPT系列、OpenVLA)在多模态任务上取得显著进展,但均基于多输入单输出(MISO)范式。我们发现该范式在多输入多输出(MIMO)场景下存在根本性局限:任务争夺共享输出通道,导致相互排斥,优化失衡,性能下降。为此,我们提出MIMO-VLA(VLASCD),一种统一训练框架,支持对话生成与决策制定的同步执行。受人类认知启发,该框架消除任务间干扰,实现高效并行处理。在CARLA自动驾驶平台上的实验表明,MIMO-VLA在MIMO设置下显著优于最先进的MISO类大语言模型、强化学习模型及视觉语言模型,为多模态多任务学习开辟新方向。

原文摘要 · Abstract (English)

Recent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input single-output (MISO) paradigm. We show that this paradigm fundamentally limits performance in multi-input multi-output (MIMO) scenarios, where parallel task execution is required. In MISO architectures, tasks compete for a shared output channel, creating mutual exclusion effects that cause unbalanced optimization and degraded performance. To address this gap, we introduce MIMO-VLA (VLASCD), a unified training framework that enables concurrent multi-task outputs, exemplified by simultaneous dialogue generation and decision-making. Inspired by human cognition, MIMO-VLA eliminates interference between tasks and supports efficient parallel processing. Experiments on the CARLA autonomous driving platform demonstrate that MIMO-VLA substantially outperforms state-of-the-art MISO-based LLMs, reinforcement learning models, and VLAs in MIMO settings, establishing a new direction for multimodal and multitask learning.

多模态决策系统并行处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。