arXiv:2508.02045cs.CL2025-08

用时间数据库自动生成问答对,让大模型时序事实回答更可信。

Harnessing Temporal Databases for Systematic Evaluation of Factual Time-Sensitive Question-Answering in Large Language Models

  • 基于时序数据库技术自动构建时序问答对
  • 提出时间准确率指标,精细评估模型时间判断能力
  • 减少人工标注依赖,适合评估特定领域时序知识

事实随时间变化,大语言模型需准确处理时序性知识。尽管时序事实问答(TSQA)任务广泛开展,现有基准常因人工瓶颈难以规模化、全面化评估。为此,我们提出 TDBench,通过利用时序数据库与数据库技术(如时序函数依赖、时序 SQL、时序连接)系统构建 TSQA 对。我们还引入新评估指标「时间准确率」,在传统答案准确率基础上,评估模型解释中时间引用的有效性,实现更细粒度的评估。对主流 LLM 的大量实验表明,TDBench 能实现可扩展、全面的 TSQA 评估,降低对人工标注的依赖,补充以 Wikipedia/Wikidata 为核心的现有评估方法,支持在应用特定数据上的模型评测。

原文摘要 · Abstract (English)

Facts change over time, making it essential for Large Language Models (LLMs) to handle time-sensitive factual knowledge accurately and reliably. Although factual Time-Sensitive Question-Answering (TSQA) tasks have been widely developed, existing benchmarks often face manual bottlenecks that limit scalable and comprehensive TSQA evaluation. To address this issue, we propose TDBench, a new benchmark that systematically constructs TSQA pairs by harnessing temporal databases and database techniques, such as temporal functional dependencies, temporal SQL, and temporal joins. We also introduce a new evaluation metric called time accuracy, which assesses the validity of time references in model explanations alongside traditional answer accuracy for a more fine-grained TSQA evaluation. Extensive experiments on contemporary LLMs show how TDBench enables scalable and comprehensive TSQA evaluation while reducing the reliance on human labor, complementing current TSQA evaluation approaches that largely center on Wikipedia/Wikidata by enabling LLM evaluation on application-specific data.

时序问答大模型评估数据库自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。