arXiv:2608.10359cs.SDcs.CL2026-08

构建首个多语言长篇语音新闻摘要与翻译基准,实现跨语言内容压缩。

VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

论文配图:VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation
图 1 · 摘自论文原文
  • 提出联合语音摘要与翻译新任务,直接生成目标语言简洁摘要。
  • 发布包含703小时语音的VoxSumm数据集,覆盖24种语言共10045对数据。
  • 发现模型在非英语目标语言上表现下降,先翻译再摘要易出错。

随着信息跨越语言边界,用户需要长篇内容的简洁跨语言表达。然而,长文档摘要研究仍以文本为中心,多语言语音研究则主要关注翻译,未注重内容压缩。为此,我们提出联合语音摘要与翻译(JSumT):从源语言长篇语音中直接生成简洁、忠实的目标语言摘要。同时,我们引入VoxSumm,首个针对该任务的多语言、跨语言基准,包含24种语言的10,045个BBC文章-摘要对,约703小时语音数据。对代表性语音-语言模型的评估显示,模型间表现差异显著:Gemini3.1-Pro一致性最高;生成英文摘要优于非英语目标语言;先完整翻译再摘要会加剧指令遵循失败。通过发布VoxSumm,我们为开发和评估能联合理解、压缩、翻译长篇语音的多语言系统奠定基础。

原文摘要 · Abstract (English)

As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual speech research has largely prioritized translation, preserving source content rather than compressing it. We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VoxSumm, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data. Our evaluation of representative speech-language models reveals pronounced variation across models and generation settings: Gemini3.1-Pro demonstrates the greatest consistency, summarization into English generally surpasses generation into non-English target languages, and translating an entire document before summarization compounds instruction-following failures. Through the release of VoxSumm, we establish a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing, and translating long-form speech.

语音摘要多语言跨语言基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。