arXiv:2602.08392cs.ROcs.AI2026-02被引 1

评测大模型在双手协作任务中的多流协同能力,发现推理强但执行差的普遍问题。

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs

  • 构建多层级基准,评估跨模态策略规划与空间对齐能力。
  • 30+模型测试显示,高层推理优秀但动作执行严重脱节。
  • 揭示感知-逻辑断联与多流干扰是核心瓶颈,适合研究具身智能者参考。

多模态大语言模型(MLLMs)虽推动了具身智能发展,但在同步双手协作任务中面临多流多模态融合的严峻挑战。本文提出ST-BiBench,一个涵盖战略协调规划、基础空间定位与细粒度动作控制的多层级评估框架。通过分析“近距悖论”——语义一致的计划却无法匹配空间视觉输入——验证模型对工作区感知与手臂选择逻辑的掌握情况。进一步探究模型能否从复杂多模态元数据直接生成高维连续动作(16维),测试30余种前沿MLLMs。结果揭示普遍存在“协调悖论”:模型在逻辑驱动的战略层面表现优异,但在多模态融合中常出现感知-逻辑脱节与多流干扰,导致精细物理执行失败。ST-BiBench为识别复杂具身任务中多流融合与跨模态对齐的关键瓶颈提供了有效平台。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have significantly advanced the landscape of embodied AI, yet transitioning to synchronized bimanual coordination introduces formidable challenges in multi-stream multimodal integration. We introduce ST-BiBench, a comprehensive multi-tier framework for evaluating spatio-temporal multimodal coordination. Our approach centers on Strategic Coordination Planning, assessing high-level cross-modal reasoning over multiple action and perception streams. To investigate the "proximity paradox"-where semantically coherent plans fail to align with spatially grounded visual inputs-we incorporate Foundational Spatial Grounding to verify workspace awareness and arm-selection logic. Furthermore, we probe model frontiers through Fine-Grained Action Control, investigating whether MLLMs can directly synthesize high-dimensional continuous action modalities (16-Dim) from complex multimodal metadata. Evaluating 30+ state-of-the-art MLLMs, we uncover a persistent and pervasive "coordination paradox"-a significant gap between high-level strategic reasoning and fine-grained physical execution. Results reveal that while frontier MLLMs excel at logic-driven strategy, they frequently suffer from perception-logic disconnection and multi-stream interference during multimodal fusion. ST-BiBench provides a platform for identifying critical bottlenecks in multi-stream multimodal fusion and cross-modal alignment for complex embodied tasks.

具身智能多模态双手协作大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。