arXiv:2603.19250cs.CL2026-03KDD

用结构化线索提升大模型在海量文档流中的表现

Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams

  • 引入结构化线索帮助模型区分不同事件
  • 在时间问答任务上性能提升最高达9.63%
  • 适合关注长文档理解与信息分离的研究者

评估语言模型在流式环境中的表现至关重要,但目前研究仍不足。现有基准要么聚焦单一复杂事件,要么为每次查询提供精心筛选的输入,未能考察多个并发事件混杂在同一文档流中时产生的冲突。我们构建了StreamBench基准,基于2016年和2025年的重大新闻故事,包含605个事件和15,354份文档,涵盖主题聚类、时间问答和摘要三个任务。为诊断模型失败原因,我们对比了有无结构化线索下的表现,发现结构化线索可使聚类任务性能提升最高达4.37%,时间问答任务提升最高达9.63%,有助于模型准确定位相关信息并分离不同事件。尽管时间推理仍是当前大模型的开放挑战,但跨任务的一致增益表明,结构化线索是未来大规模文档流研究的重要方向。

原文摘要 · Abstract (English)

Evaluating language models in streaming environments is critical, yet underexplored. Existing benchmarks either focus on single complex events or provide curated inputs for each query, and do not evaluate models under the conflicts that arise when multiple concurrent events are mixed within the same document stream. We introduce StreamBench, a benchmark built from major news stories in 2016 and 2025, comprising 605 events and 15,354 documents across three tasks: Topic Clustering, Temporal Question Answering, and Summarization. To diagnose how models fail, we compare performance with and without structural cues, which organize key facts by event. We find that structural cues improve performance on clustering (up to +4.37%) and temporal QA (up to +9.63%), helping models locate relevant information and separate distinct events. While temporal reasoning remains an open challenge inherent to current LLMs, consistent gains across tasks show that structural cues are a promising direction for future work in massive document streams.

文档流结构线索大模型评估时间问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。