用自动合成数据训练视觉语言模型,提升多模态工具推理能力。
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- 构建28.5K任务的多模态轨迹数据集M-TRACE,支持模仿学习
- 通过11K自动生成偏好对,实现分步偏好优化,性能超越主流模型
- 适合需要强工具调用能力的多模态智能体研究者使用
视觉语言模型(VLMs)越来越多地作为控制器,通过访问外部工具进行复杂推理与决策,但其效果受限于高质量多模态轨迹数据稀缺及人工标注成本高。我们提出一种以视觉为中心的智能体微调框架,可自动合成多模态轨迹,生成分步偏好对,并训练出具备鲁棒工具使用推理能力的VLM控制器。该流程首先构建了大规模数据集M-TRACE,包含28.5K个多模态任务和177K条经验证的轨迹,支持基于模仿的学习。在此基础上,我们开发了在M-TRACE上微调的MATRIX Agent,用于分步工具推理。为进一步提升对齐精度,我们引入11K条自动生成的偏好对Pref-X,通过分步偏好学习优化MATRIX。在Agent-X、GTA和GAIA三个基准测试中,MATRIX持续优于开源与闭源的VLM,展现出可扩展且高效的多模态工具使用能力。数据与代码已公开于https://github.com/mbzuai-oryx/MATRIX。
原文摘要 · Abstract (English)
Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories and the cost of manual annotation. We address this challenge with a vision-centric agent tuning framework that automatically synthesizes multimodal trajectories, generates step-wise preference pairs, and trains a VLM controller for robust tool-use reasoning. Our pipeline first constructs M-TRACE, a large-scale dataset of 28.5K multimodal tasks with 177K verified trajectories, enabling imitation-based trajectory tuning. Building on this, we develop MATRIX Agent, a controller finetuned on M-TRACE for step-wise tool reasoning. To achieve finer alignment, we further introduce Pref-X, a set of 11K automatically generated preference pairs, and optimize MATRIX on it via step-wise preference learning. Across three benchmarks, Agent-X, GTA, and GAIA, MATRIX consistently surpasses both open- and closed-source VLMs, demonstrating scalable and effective multimodal tool use. Our data and code is avaliable at https://github.com/mbzuai-oryx/MATRIX.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。