arXiv:2606.03005cs.CVcs.AI2026-06被引 1

用智能执行框架提升冻结模型的多模态能力,无需重训练

MUSE: A Unified Agentic Harness for MLLMs

论文配图:MUSE: A Unified Agentic Harness for MLLMs
图 1 · 摘自论文原文
  • 构建可组合模块框架,增强模型的任务表示与视觉处理能力
  • 在多个基准上实现显著提升,复杂任务性能增幅最大
  • 验证器引导修复机制可解决多数失败问题,适合追求效率的研究者

尽管多模态大语言模型(MLLM)进展迅速,但在人类轻松完成的任务上仍表现不佳,例如从截图导航网格迷宫或选择正确拼图块。我们提出一个互补问题:在不重新训练模型的前提下,仅通过改进其执行框架,能激发多少潜能?为此,我们提出MUSE,一种统一的多模态结构化执行框架,可无缝集成任意现成的MLLM,包含任务表征、视觉处理、感知工具使用、结构化解析、确定性验证及验证器引导修复等可组合模块,无需模型重训。我们在涵盖视觉空间规划、视觉感知、多模态推理和细粒度视觉辨别等多个基准上评估,使用多种先进MLLM。MUSE在所有设置中均带来一致性能提升,尤其在高难度任务上增益显著。进一步分析表明,多数MLLM失败源于执行框架缺陷而非模型本身不足,可通过验证器引导修复解决,无需改动模型。这些发现强调了代理式多模态框架作为关键但被忽视的设计维度,为超越模型中心优化提供了新路径。

原文摘要 · Abstract (English)

Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze from a screenshot or selecting the correct puzzle piece. Rather than retraining the model, we ask a complementary question: how much capability can be elicited from a frozen MLLM purely by improving the execution scaffold around it? We introduce MUSE, a multimodal unified structured execution harness that wraps any off-the-shelf MLLM with composable modules for task representation, visual processing, perception tool use, structured parsing, deterministic verification, and verifier-guided repair, without any model retraining. We evaluate MUSE across diverse benchmarks spanning visual spatial planning, visual perception, multimodal reasoning, and fine-grained visual discrimination, using multiple state-of-the-art MLLMs. MUSE delivers consistent gains over the bare model in all settings, with the largest jumps on challenging instances. Further analysis reveals that many MLLM failures arise from harness-level shortcomings rather than fundamental model deficits, and can be addressed through verifier-guided repair without touching the model. These findings highlight the agentic multimodal harness as a critical yet underexplored design dimension, offering an orthogonal avenue for improving MLLMs beyond model-centric optimization.

多模态执行框架智能体推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。