arXiv:2510.11454cs.SDcs.AI2025-10被引 5

让音频大模型调用外部工具,提升推理准确性和可解释性

Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning

  • 引入外部工具协同推理,实现对音频信号的分步分析与处理
  • 多模型测试显示平均准确率提升4-5个百分点,最高达72.1%
  • 适合需要精准音频分析的科研与工业场景,如医疗听觉诊断

近期大型多模态模型在音频理解方面表现出强大能力,但多数系统仅依赖端到端推理,限制了在需结构化知识或专业信号分析任务中的可解释性与准确性。本文提出Audio-Maestro——一种工具增强的音频推理框架,使音频语言模型能够自主调用外部工具,并将带时间戳的输出结果融入推理流程。该设计使模型通过专用工具分析、转换和解释音频信号,而非仅依赖端到端推断。实验表明,Audio-Maestro持续提升通用音频推理性能:Gemini-2.5-flash在MMAU-Test上的平均准确率从67.4%提升至72.1%,DeSTA-2.5从58.3%升至62.8%,GPT-4o从60.8%增至63.9%。据我们所知,Audio-Maestro是首个将结构化工具输出整合进大音频语言模型推理过程的框架。

原文摘要 · Abstract (English)

Recent advancements in large multimodal models (LMMs) have shown strong capabilities in audio understanding. However, most systems rely solely on end-to-end reasoning, limiting interpretability and accuracy for tasks that require structured knowledge or specialized signal analysis. In this work, we present Audio-Maestro -- a tool-augmented audio reasoning framework that enables audio-language models to autonomously call external tools and integrate their timestamped outputs into the reasoning process. This design allows the model to analyze, transform, and interpret audio signals through specialized tools rather than relying solely on end-to-end inference. Experiments show that Audio-Maestro consistently improves general audio reasoning performance: Gemini-2.5-flash's average accuracy on MMAU-Test rises from 67.4% to 72.1%, DeSTA-2.5 from 58.3% to 62.8%, and GPT-4o from 60.8% to 63.9%. To our knowledge, Audio-Maestro is the first framework to integrate structured tool output into the large audio language model reasoning process.

音频理解工具增强多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。