arXiv:2510.00161cs.CL2025-10

TAMA让智能体用工具看懂操作视频,提升理解准确率。

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

  • 不依赖训练,通过调用多媒体工具实现多模态交叉推理。
  • 在ProMQA-Assembly数据集上显著提升GPT-5和MiMo-VL性能。
  • 适合开发厨房、制造等场景的操作辅助系统。

操作性活动助手可在日常生活(如烹饪、组装家具)和专业场景(如制造、生物实验)中辅助人类。尽管应用前景广阔,针对此类助手的系统开发仍处于探索阶段。本文提出一种新框架TAMA(Tool-Augmented Multimodal Agent),用于程序化活动理解。TAMA在无需训练的条件下,通过调用多媒体返回工具,实现多模态信息的交错推理。在多模态程序问答数据集ProMQA-Assembly上的实验表明,该方法显著提升了视觉语言模型(尤其是GPT-5和MiMo-VL)的性能。消融实验进一步验证了两大核心特性:多媒体返回工具与代理式灵活工具选择的有效性。我们认为,本框架及实验结果推动了‘以图像思考’范式在视频与多模态任务中的发展,也为操作性活动助手的构建提供了支持。

原文摘要 · Abstract (English)

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological experiments. Despite its potential use cases, the system development tailored for such an assistant is still underexplored. In this paper, we propose a novel framework, called TAMA, a Tool-Augmented Multimodal Agent, for procedural activity understanding. TAMA enables interleaved multimodal reasoning by making use of multimedia-returning tools in a training-free setting. Our experimental result on the multimodal procedural QA dataset, ProMQA-Assembly, shows that our approach can improve the performance of vision-language models, especially GPT-5 and MiMo-VL. Furthermore, our ablation studies provide empirical support for the effectiveness of two features that characterize our framework, multimedia-returning tools and agentic flexible tool selection. We believe our proposed framework and experimental results facilitate the thinking with images paradigm for video and multimodal tasks, let alone the development of procedural activity assistants.

多模态智能体操作理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。