首份系统性评估大音频语言模型的综述,构建四维评价体系。
Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
- 按四大目标维度分类评测任务:听觉感知、知识推理、对话能力、公平安全。
- 梳理现有基准,揭示评估碎片化问题,提出统一框架。
- 适合研究者、开发者参考,助力音频语言模型健康发展。
随着大型音频语言模型(LALMs)的发展,这些模型通过引入听觉能力增强了大语言模型(LLMs)的通用性。尽管已有众多基准用于评估LALMs性能,但它们仍呈碎片化状态,缺乏系统性分类。为弥合这一差距,本文开展全面综述,提出一套基于目标的系统性LALM评估分类体系,涵盖四个维度:(1) 通用听觉感知与处理,(2) 知识与推理,(3) 对话能力,(4) 公平性、安全性与可信度。本文对各维度进行详尽分析,指出当前挑战,并展望未来方向。据我们所知,这是首个专注于LALM评估的综述,为社区提供明确指导。我们将持续维护所调研论文集合,以支持该领域持续发展。
原文摘要 · Abstract (English)
With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous benchmarks have emerged to assess LALMs' performance, they remain fragmented and lack a structured taxonomy. To bridge this gap, we conduct a comprehensive survey and propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. We provide detailed overviews within each category and highlight challenges in this field, offering insights into promising future directions. To the best of our knowledge, this is the first survey specifically focused on the evaluations of LALMs, providing clear guidelines for the community. We will release the collection of the surveyed papers and actively maintain it to support ongoing advancements in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。