arXiv:2506.21783cs.CLcs.AI2025-06中稿 · ICTIR 2025 co-loca…被引 4

提出新基准TLQA,评估大模型列表生成与时间理解能力

Evaluating List Construction and Temporal Understanding capabilities of Large Language Models

  • 设计需同时处理列表构建与时间对齐的问答任务
  • 现有模型在封闭域下漏答率高,开放域需依赖检索增强
  • 适合研究大模型时序推理与结构化输出的学者使用

大型语言模型在众多自然语言任务中表现出色,但在涉及多实体的时间理解任务中仍易产生幻觉和错误。这些模型常无法准确关联实体与时间区间,未能完整生成实体列表,或无法对特定时间边界内的事件进行推理。现有工作未充分评估模型在列表答案构造场景下的隐式与显式时间理解能力。为此,我们提出了基于时间参考的列表问答基准TLQA,要求答案以与时间周期对齐的列表形式呈现。该基准首次同时考察列表构建与时间理解能力,目前尚无类似研究。我们在闭源与开源设置下评估了先进生成模型在TLQA上的表现,结果表明当前模型存在显著缺陷:在闭源设置中难以提供完整答案且事实时间对齐能力差;在开源设置中则依赖外部检索。这为未来研究指明了方向。数据集与代码已开源。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated immense advances in a wide range of natural language tasks. However, these models are susceptible to hallucinations and errors on particularly temporal understanding tasks involving multiple entities in answers. In such tasks, they fail to associate entities with accurate time intervals, generate a complete list of entities in answers or reason about events associated with specific temporal bounds. Existing works do not extensively evaluate the abilities of the model to perform implicit and explicit temporal understanding in a list answer construction setup. To bridge this gap, we propose the Time referenced List based Question Answering or TLQA benchmark that requires structured answers in list format aligned with corresponding time periods. Our TLQA benchmark, requires both list construction and temporal understanding simultaneously, which to the best of our knowledge has not been explored in prior benchmarks. We investigate the temporal understanding and list construction capabilities of state-of-the-art generative models on TLQA in closed-book and open-domain settings. Our findings reveal significant shortcomings in current models, particularly their inability to provide complete answers and temporally align facts in a closed-book setup and the need to improve retrieval in open-domain setup, providing clear future directions for research on TLQA. The benchmark and code at https://github.com/elixir-research-group/TLQA.

大模型评测时间理解列表生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。