arXiv:2508.12282cs.CLcs.IR2025-08被引 3

构建中文时间敏感问答数据集,评估RAG系统时序推理能力

A Question Answering Dataset for Temporal-Sensitive Retrieval-Augmented Generation

  • 从2019-2024年超30万篇新闻构建,覆盖绝对、相对、聚合三类时间问题
  • 含5176个高质量问题,支持单/多文档场景,验证时序对齐与逻辑一致性
  • 经规则、大模型和人工三重验证,适合作为时间敏感RAG系统的基准

我们提出ChronoQA,一个大规模中文问答数据集,专门用于评估检索增强生成(RAG)系统中的时序推理能力。该数据集基于2019至2024年间超过30万篇新闻文章构建,包含5,176个高质量问题,涵盖绝对、聚合和相对时间类型,且包含显式与隐式时间表达。数据集支持单文档与多文档场景,反映真实世界中对时序对齐与逻辑一致性的需求。ChronoQA具备全面的结构化标注,并经过规则、大语言模型及人工三阶段验证,确保数据质量。该数据集提供动态、可靠且可扩展的评估资源,支持广泛的时间任务评测,是推进时间敏感检索增强问答系统的重要基准。

原文摘要 · Abstract (English)

We introduce ChronoQA, a large-scale benchmark dataset for Chinese question answering, specifically designed to evaluate temporal reasoning in Retrieval-Augmented Generation (RAG) systems. ChronoQA is constructed from over 300,000 news articles published between 2019 and 2024, and contains 5,176 high-quality questions covering absolute, aggregate, and relative temporal types with both explicit and implicit time expressions. The dataset supports both single- and multi-document scenarios, reflecting the real-world requirements for temporal alignment and logical consistency. ChronoQA features comprehensive structural annotations and has undergone multi-stage validation, including rule-based, LLM-based, and human evaluation, to ensure data quality. By providing a dynamic, reliable, and scalable resource, ChronoQA enables structured evaluation across a wide range of temporal tasks, and serves as a robust benchmark for advancing time-sensitive retrieval-augmented question answering systems.

问答系统时序推理RAG中文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。