arXiv:2508.15274cs.CLcs.IR2025-08

用大模型自动挖掘事件时间常识,构建新数据集提升语言模型推理能力

TComQA: Extracting Temporal Commonsense from Text

  • 利用大模型从文本中自动提取事件持续时间等时间常识
  • 构建的TComQA数据集经众包验证,提取精度超80%
  • 该数据集可有效提升模型在时间推理任务上的表现

理解事件需要把握其时间背景,但自然语言中常未明确提及。例如,机器难以推断博物馆参观可能持续数小时而非数月。研究表明,即使先进大语言模型在涉及时间常识推理的任务上仍表现不佳,因其在文本中极少被显式描述。为此,本文研究大模型提取文本中时间常识的能力,并评估多种实验设置的效果。提出一种基于大模型的时间常识提取流程,从SAMSum和RealNews语料中自动挖掘并构建TComQA数据集。该数据集经众包验证,时间常识提取精度超过80%。使用TComQA训练的模型在时间问答任务上优于在现有数据集上微调的大模型。

原文摘要 · Abstract (English)

Understanding events necessitates grasping their temporal context, which is often not explicitly stated in natural language. For example, it is not a trivial task for a machine to infer that a museum tour may last for a few hours, but can not take months. Recent studies indicate that even advanced large language models (LLMs) struggle in generating text that require reasoning with temporal commonsense due to its infrequent explicit mention in text. Therefore, automatically mining temporal commonsense for events enables the creation of robust language models. In this work, we investigate the capacity of LLMs to extract temporal commonsense from text and evaluate multiple experimental setups to assess their effectiveness. Here, we propose a temporal commonsense extraction pipeline that leverages LLMs to automatically mine temporal commonsense and use it to construct TComQA, a dataset derived from SAMSum and RealNews corpora. TComQA has been validated through crowdsourcing and achieves over 80\% precision in extracting temporal commonsense. The model trained with TComQA also outperforms an LLM fine-tuned on existing dataset of temporal question answering task.

时间推理常识挖掘大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。