构建时间推理基准,揭示大模型在日期理解中的两大偏差
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
- 设计包含190个问题的基准,覆盖多种日期格式与推理类型
- 发现表示层与逻辑层存在显著时间推理偏差
- 适合关注时序建模与语言模型公平性的研究者
本文提出 DateLogicQA,一个包含190个问题的基准,涵盖多样的日期格式、时间上下文和推理类型。我们引入语义完整性度量来评估分词质量,并分析两种偏差:影响嵌入表示的表示层偏差,以及影响推理输出的逻辑层偏差。研究结果全面评估了大语言模型在时间推理方面的能力与局限,凸显了其准确处理时间数据的关键挑战。
原文摘要 · Abstract (English)
This paper introduces DateLogicQA, a benchmark with 190 questions covering diverse date formats, temporal contexts, and reasoning types. We propose the Semantic Integrity Metric to assess tokenization quality and analyse two biases: Representation-Level Bias, affecting embeddings, and Logical-Level Bias, influencing reasoning outputs. Our findings provide a comprehensive evaluation of LLMs' capabilities and limitations in temporal reasoning, highlighting key challenges in handling temporal data accurately.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。