arXiv:2601.13539cs.SD2026-01被引 6

构建超长语音评测基准,推动模型跨任务理解能力发展。

LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech

  • 从多源数据构建百万级10分钟长语音片段,支持多任务标注。
  • 现有模型在长语音上表现差,多任务间存在性能权衡。
  • 适合研究语音理解、多模态推理与长序列建模的学者使用。

近年来,音频-语言模型在短时语音任务中取得显著进展。然而,会议转录、口语文档理解及对话分析等真实场景需要模型具备处理和推理长时音频的能力。本文提出 LongSpeech,一个大规模、可扩展的长语音评测基准,用于评估和推进语音模型在长音频上的表现。该基准包含超过10万段约10分钟长的语音片段,涵盖语音识别(ASR)、语音翻译、摘要生成、语言检测、说话人计数、内容分离和问答等多种丰富标注。我们提出一个可复现的构建流程,可从多样来源生成长语音基准,支持未来扩展。初步实验显示,当前先进模型在长语音任务中存在显著性能差距,往往在单一任务上表现突出却牺牲其他任务,且难以进行高层推理。这些发现凸显了本基准的挑战性。该基准将向研究社区公开。

原文摘要 · Abstract (English)

Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis require robust models capable of processing and reasoning over long-form audio. In this work, we present LongSpeech, a large-scale and scalable benchmark specifically designed to evaluate and advance the capabilities of speech models on long-duration audio. LongSpeech comprises over 100,000 speech segments, each approximately 10 minutes long, with rich annotations for ASR, speech translation, summarization, language detection, speaker counting, content separation, and question answering. We introduce a reproducible pipeline for constructing long-form speech benchmarks from diverse sources, enabling future extensions. Our initial experiments with state-of-the-art models reveal significant performance gaps, with models often specializing in one task at the expense of others and struggling with higher-level reasoning. These findings underscore the challenging nature of our benchmark. Our benchmark will be made publicly available to the research community.

长语音多任务评测基准语音理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。