arXiv:2601.19673cs.SDcs.AI2026-01Conference of the …被引 1

构建新基准,评估多模态大模型的音频推理能力

A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models

  • 提出跨任务推理的音频评测框架ART
  • 首次测试模型融合不同音频任务的能力
  • 适合研究多模态推理与音频理解的学者

现有针对多模态大模型音频模态的评测基准主要孤立地测试语音任务,如说话人分离或性别识别。这些方法无法验证模型是否具备将不同类别音频任务结合进行推理的能力。为解决此问题,我们提出音频推理任务(Audio Reasoning Tasks, ART),一个新基准,用于评估多模态模型在处理需对音频信号进行综合推理的问题时的表现。

原文摘要 · Abstract (English)

The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identification in isolation. Whether a multimodal model can answer the questions that require reasoning skills to combine audio tasks of different categories, cannot be verified with their use. To address this issue, we propose Audio Reasoning Tasks (ART), a new benchmark for assessing the ability of multimodal models to solve problems that require reasoning over audio signal.

音频推理多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。