评测AI终端代理处理音视频文件任务的能力,填补真实工作流评估空白。
MMTB: Evaluating Terminal Agents on Multimedia-File Tasks

- 构建音视频任务基准MMTB,涵盖105项跨5类的多媒体操作任务
- 提出支持音视频感知的Terminus-MM框架,提升代理对多媒体证据的理解能力
- 揭示多媒体信息如何影响任务执行路径,适合多模态智能体研究者使用
终端为人工智能代理提供了强大接口,可调用多样工具自动化复杂流程,但现有终端代理评估基准主要聚焦于文本、代码和结构化文件任务。然而,许多真实工作流需要直接处理音频和视频文件。此类任务要求代理不仅理解多媒体内容,还需将听觉与视觉信息转化为恰当操作。为此,我们提出多模态终端基准MMTB,包含105项任务,覆盖5个元类别,支持代理直接操作音视频文件。同时,我们设计了Terminus-MM,基于Terminus-KIRA扩展音频与视频感知能力。二者共同支持对多媒体终端代理的可控研究,揭示不同多媒体访问形式如何影响任务结果,并决定代理依赖何种证据构建可执行终端工作流。MMTB媒体与元数据已公开发布于https://huggingface.co/datasets/mm-tbench/mmtb-media。
原文摘要 · Abstract (English)
Terminals provide a powerful interface for AI agents by exposing diverse tools for automating complex workflows, yet existing terminal-agent benchmarks largely focus on tasks grounded in text, code, and structured files. However, many real-world workflows require practitioners to work directly with audio and video files. Working with such multimedia files calls for terminal agents not only to understand multimedia content, but also to convert auditory and visual evidence across related files into appropriate actions. To evaluate terminal agents on multimedia-file tasks, we introduce MultiMedia-TerminalBench (MMTB), a benchmark of 105 tasks across 5 meta-categories where terminal agents directly operate with audio and video files. Alongside MMTB, we propose Terminus-MM, a multimedia harness that extends Terminus-KIRA with audio and video perception for terminal agents. Together, MMTB and Terminus-MM support a controlled study of multimedia terminal agents, revealing how different forms of multimedia access shape task outcomes and determine which evidence agents rely on to construct executable terminal workflows. MMTB media and metadata are released at https://huggingface.co/datasets/mm-tbench/mmtb-media
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。