构建119小时古文语音数据集,推动多模态模型在古典文学语音任务上的研究。
MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus
- 构建涵盖六类任务的古文语音多任务数据集
- 10个大模型在测试集上表现仍不理想,挑战显著
- 提供专用评估指标,适合古文语音与多模态研究者使用
随着多模态大语言模型(MLLMs)的快速发展,其在古代汉语研究中的潜力备受关注。现有研究主要聚焦于文本与视觉模态,而该领域中的语音语料库仍严重缺乏。为填补这一空白,我们提出了多任务古典文学语音语料库(MCGA),包含119小时、22,000个音频样本,覆盖六项任务:自动语音识别(ASR)、语音到文本翻译(S2TT)、语音情感描述(SEC)、口语问答(SQA)、语音理解(SU)和语音推理(SR)。通过对十种MLLM的评估,实验结果表明当前模型在MCGA测试集上仍面临重大挑战。此外,我们引入了针对SEC的领域专用指标,以及衡量语音与文本能力一致性的评估方法。MCGA已公开发布,以促进更鲁棒的MLLM发展。
原文摘要 · Abstract (English)
With the rapid advancement of Multimodal Large Language Models (MLLMs), their potential has gained significant attention in Chinese Classical Studies (CCS). While existing research primarily focuses on text and visual modalities, the audio corpus within this domain remains largely underexplored. To bridge this gap, we introduce the Multi-task Classical Chinese Literary Genre Audio Corpus (MCGA), a 119-hour corpus comprising 22,000 audio samples. It encompasses a diverse range of literary genres across six tasks: Automatic Speech Recognition (ASR), Speech-to-Text Translation (S2TT), Speech Emotion Captioning (SEC), Spoken Question Answering (SQA), Speech Understanding (SU), and Speech Reasoning (SR). Through the evaluation of ten MLLMs, our experimental results demonstrate that current MLLMs still face substantial challenges on the MCGA test set. Furthermore, we introduce a domain-specific metric for SEC and a metric to measure the consistency between speech and text capabilities. We release MCGA to the public to facilitate the development of more robust MLLMs. MCGA Corpus: https://github.com/yxduir/MCGA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。