arXiv:2511.01101cs.CL2025-11EMNLP被引 2

构建首个面向时间序列事实验证的高质量基准数据集

TSVer: A Benchmark for Fact Verification Against Time-Series Evidence

  • 基于真实世界声明与400条时间序列,设计多步标注流程
  • 模型在事实判断上准确率仅63.57%,论据生成能力不足
  • 适合研究时序推理、可信AI与自动验证系统的学者

对时间与数值数据(如时间序列)进行推理是事实核查的关键。尽管近期已开发多种系统处理此类证据,但现有数据集仍受限于缺乏结构化证据、结论理由不足或依赖合成声明。本文提出TSVer,一个专注于时间序列证据的事实验证新基准。该数据集包含来自41家事实核查机构的304个真实声明,以及覆盖多元领域的400条时间序列。每条声明均标注相关的时间范围,并附有判决结果及解释,说明证据如何支持结论。通过大模型辅助的多步骤标注流程,实现判决一致率kappa=0.77。我们还建立了基线模型,结果显示即使是Gemini-2.5-Pro等前沿模型,在判决准确率上也仅为63.57%,论据生成评估得分(Ev2R)为47.36,表明当前模型仍难以有效处理时间序列推理。

原文摘要 · Abstract (English)

Reasoning over temporal and numerical data, such as time series, is a crucial aspect of fact-checking. While many systems have recently been developed to handle this form of evidence, their evaluation remains limited by existing datasets, which often lack structured evidence, provide insufficient justifications for verdicts, or rely on synthetic claims. In this paper, we introduce TSVer, a new benchmark dataset for fact verification focusing on temporal and numerical reasoning with time-series evidence. TSVer contains 304 real-world claims sourced from 41 fact-checking organizations and a curated database of 400 time series covering diverse domains. Each claim is annotated with time frames across all pertinent time series, along with a verdict and justifications reflecting how the evidence is used to reach the verdict. Using an LLM-assisted multi-step annotation process, we improve the quality of our annotations and achieve an inter-annotator agreement of kappa=0.77 on verdicts. We also develop a baseline for verifying claims against time-series evidence and show that even the state-of-the-art reasoning models like Gemini-2.5-Pro are challenged by time series, achieving a 63.57 accuracy score on verdicts and an Ev2R score of 47.36 on verdict justifications.

事实核查时间序列基准测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。