MindWatcher能自主调用工具并多模态推理,解决复杂现实问题。
MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- 采用思维与工具调用交替的机制,灵活决策何时何地用工具。
- 在多模态任务中表现超越大模型,本地图像库覆盖8类物体。
- 适合需要跨模态推理与工具协同的智能体研发人员使用。
传统基于工作流的智能体在需要调用外部工具的现实问题上表现出智能局限。能够自主推理与工具调用的工具集成推理(TIR)代理正成为处理多步交互复杂决策任务的强大方法。本文提出MindWatcher,一种融合交错思维与多模态思维链(CoT)推理的TIR代理。它可自主决定是否及如何调用多种工具并协调其使用,无需人类提示或预设流程。交错思维范式允许模型在任意中间阶段切换思考与工具调用;多模态CoT能力支持在推理过程中操作图像,以获得更精准的搜索结果。我们构建了自动化数据审计与评估流水线,并结合人工标注的高质量数据集进行训练,同时设计名为MindWatcher-Evaluate Bench(MWE-Bench)的基准测试框架。MindWatcher配备全面的辅助推理工具,可应对广泛领域的多模态问题。一个涵盖汽车、动物、植物等八类对象的大规模本地图像检索数据库,使其在小模型规模下仍具备强健的物体识别能力。最后,我们设计了更高效的训练基础设施,显著提升训练速度与硬件利用率。实验表明,通过优越的工具调用能力,MindWatcher在性能上达到甚至超过更大或更新的模型,并揭示了智能体训练中的关键洞见,如智能体强化学习中的遗传继承现象。
原文摘要 · Abstract (English)
Traditional workflow-based agents exhibit limited intelligence when addressing real-world problems requiring tool invocation. Tool-integrated reasoning (TIR) agents capable of autonomous reasoning and tool invocation are rapidly emerging as a powerful approach for complex decision-making tasks involving multi-step interactions with external environments. In this work, we introduce MindWatcher, a TIR agent integrating interleaved thinking and multimodal chain-of-thought (CoT) reasoning. MindWatcher can autonomously decide whether and how to invoke diverse tools and coordinate their use, without relying on human prompts or workflows. The interleaved thinking paradigm enables the model to switch between thinking and tool calling at any intermediate stage, while its multimodal CoT capability allows manipulation of images during reasoning to yield more precise search results. We implement automated data auditing and evaluation pipelines, complemented by manually curated high-quality datasets for training, and we construct a benchmark, called MindWatcher-Evaluate Bench (MWE-Bench), to evaluate its performance. MindWatcher is equipped with a comprehensive suite of auxiliary reasoning tools, enabling it to address broad-domain multimodal problems. A large-scale, high-quality local image retrieval database, covering eight categories including cars, animals, and plants, endows model with robust object recognition despite its small size. Finally, we design a more efficient training infrastructure for MindWatcher, enhancing training speed and hardware utilization. Experiments not only demonstrate that MindWatcher matches or exceeds the performance of larger or more recent models through superior tool invocation, but also uncover critical insights for agent training, such as the genetic inheritance phenomenon in agentic RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。