构建首个评估模糊、隐含与明确时间指代的基准,揭示大模型在复杂时间推理中的短板。
TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References
- 设计问答式合成数据集,覆盖三类时间指代:明确、相对说话时间、模糊
- 4个主流大模型在事件数多或指代模糊时性能显著下降,模糊类准确率最低
- 基于人类调研生成模糊问题真值,适合评估时间推理能力的研究者使用
理解并解决时间指代对自然语言理解至关重要,因为日常交流中常涉及过去或未来。尽管现有基准已评估系统处理时间指代的能力,但对特定类型时间指代的系统性评估仍有限。为此,我们提出TRAVELER,一个新型合成基准数据集,采用问答范式,包含涉及时间指代的问题及其正确答案。TRAVELER评估模型对明确、相对于说话时间的隐含、以及模糊时间指代的解析能力。除考察不同时间指代类型下先进大模型的表现外,该基准还支持评估事件集合长度对性能的影响。对于模糊时间指代类别,真值通过Prolific平台的人类调研确定,方法参照Kenneweg等人的流程。为展示基准适用性,我们使用涵盖3,300个问题的问答任务评估了四个最先进的大模型。结果表明,尽管这些模型在事件少且指代明确时表现良好,但随着事件数量增加和指代变得不明确,性能明显下降。尤其在模糊问题类别中,所有模型表现最差。
原文摘要 · Abstract (English)
Understanding and resolving temporal references is essential in Natural Language Understanding as we often refer to the past or future in daily communication. Although existing benchmarks address a system's ability to reason about and resolve temporal references, systematic evaluation of specific temporal references remains limited. Towards closing this gap, we introduce TRAVELER, a novel synthetic benchmark dataset that follows a Question Answering paradigm and consists of questions involving temporal references with the corresponding correct answers. TRAVELER assesses models' abilities to resolve explicit, implicit relative to speech time, and vague temporal references. Beyond investigating the performance of state-of-the-art LLMs depending on the type of temporal reference, our benchmark also allows evaluation of performance in relation to the length of the set of events. For the category of vague temporal references, ground-truth answers were established via human surveys on Prolific, following a procedure similar to the one from Kenneweg et al. To demonstrate the benchmark's applicability, we evaluate four state-of-the-art LLMs using a question-answering task encompassing 3,300 questions. Our findings show that while the benchmarked LLMs can answer questions over event sets with a handful of events and explicit temporal references successfully, performance clearly deteriorates with larger event set length and when temporal references get less explicit. Notably, the vague question category exhibits the lowest performance across all models. The benchmark is publicly available at: https://gitlab.ub.uni-bielefeld.de/s.kenneweg/TRAVELER
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。