首个面向金融场景的音频大模型评测基准,填补领域空白。
FinAudio: A Benchmark for Audio Large Language Models in Financial Applications
- 构建三类金融音频任务:短/长语音识别与长音频摘要。
- 涵盖4个数据集,含首创的金融音频摘要数据集。
- 评测7个主流音频大模型,揭示其在金融场景中的不足。
音频大语言模型(AudioLLMs)在对话、音频理解及自动语音识别(ASR)等任务上表现优异,但在金融领域尚无专用评测基准。本文提出首个面向金融应用的AudioLLM评测基准 extsc{FinAudio},针对金融音频特性设计三类任务:短音频ASR、长音频ASR和长音频摘要。我们构建了两个短音频与两个长音频数据集,并开发首个金融音频摘要数据集,形成完整 extsc{FinAudio} 基准。在此基础上,评估7个主流AudioLLMs的表现,揭示其在金融场景下的局限性,为模型优化提供方向。所有数据集与代码将公开。
原文摘要 · Abstract (English)
Audio Large Language Models (AudioLLMs) have received widespread attention and have significantly improved performance on audio tasks such as conversation, audio understanding, and automatic speech recognition (ASR). Despite these advancements, there is an absence of a benchmark for assessing AudioLLMs in financial scenarios, where audio data, such as earnings conference calls and CEO speeches, are crucial resources for financial analysis and investment decisions. In this paper, we introduce \textsc{FinAudio}, the first benchmark designed to evaluate the capacity of AudioLLMs in the financial domain. We first define three tasks based on the unique characteristics of the financial domain: 1) ASR for short financial audio, 2) ASR for long financial audio, and 3) summarization of long financial audio. Then, we curate two short and two long audio datasets, respectively, and develop a novel dataset for financial audio summarization, comprising the \textsc{FinAudio} benchmark. Then, we evaluate seven prevalent AudioLLMs on \textsc{FinAudio}. Our evaluation reveals the limitations of existing AudioLLMs in the financial domain and offers insights for improving AudioLLMs. All datasets and codes will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。