arXiv:2504.17366cs.CLcs.AI2025-04ACL被引 2

首个面向直播口语长文本的基准测试,揭示当前模型在冗余语境下的理解短板。

LiveLongBench: Tackling Long-Context Understanding for Spoken Texts from Live Streams

  • 构建来自直播流的真实口语长文本数据集,涵盖检索、推理与混合任务
  • 现有模型在高冗余输入下表现差,无方法在所有任务中稳定领先
  • 提出新基线模型,有效应对口语冗余,适合电商等真实场景应用

长上下文理解在自然语言处理中面临巨大挑战,尤其针对以语音为基础、高冗余和信息密度不均的现实对话。尽管大语言模型(LLMs)在现有基准上表现优异,但这些数据集未能反映真实文本的复杂性,限制了其实际应用。为此,我们构建了首个源自直播流的口语长文本数据集,真实呈现冗余丰富和对话式特征。设计三类任务:依赖检索、依赖推理及混合任务。评估主流LLMs与专用方法在这些任务中的表现。结果表明,当前方法存在明显任务偏好,在高冗余输入下性能显著下降,且无单一方法始终领先。我们提出一种新基线,更有效处理口语冗余,在各类任务中均取得良好表现。研究揭示了当前方法的关键局限,并为未来改进提供方向。该基准填补了长上下文口语理解评估的空白,为开发真实场景电商系统提供实用基础。代码与数据集已公开于https://github.com/Yarayx/livelongbench。

原文摘要 · Abstract (English)

Long-context understanding poses significant challenges in natural language processing, particularly for real-world dialogues characterized by speech-based elements, high redundancy, and uneven information density. Although large language models (LLMs) achieve impressive results on existing benchmarks, these datasets fail to reflect the complexities of such texts, limiting their applicability to practical scenarios. To bridge this gap, we construct the first spoken long-text dataset, derived from live streams, designed to reflect the redundancy-rich and conversational nature of real-world scenarios. We construct tasks in three categories: retrieval-dependent, reasoning-dependent, and hybrid. We then evaluate both popular LLMs and specialized methods to assess their ability to understand long-contexts in these tasks. Our results show that current methods exhibit strong task-specific preferences and perform poorly on highly redundant inputs, with no single method consistently outperforming others. We propose a new baseline that better handles redundancy in spoken text and achieves strong performance across tasks. Our findings highlight key limitations of current methods and suggest future directions for improving long-context understanding. Finally, our benchmark fills a gap in evaluating long-context spoken language understanding and provides a practical foundation for developing real-world e-commerce systems. The code and benchmark are available at https://github.com/Yarayx/livelongbench.

长文本理解口语处理直播分析大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。