arXiv:2409.14065cs.CLcs.LG2024-09EMNLP被引 4

提出时间一致事实性探测任务,提升大模型问答一致性。

Temporally Consistent Factuality Probing for Large Language Models

  • 构建前缀式英文查询数据集,支持时序一致性评估。
  • 多数大模型在时间一致性任务上表现不佳。
  • 提出新训练框架,显著改善模型时序事实一致性。

大语言模型作为替代知识库的广泛应用,要求其具备事实正确性与一致性,尤其在改写查询时。现有评测数据集和指标多基于主语-关系-宾语的简单结构,且依赖当前关联,限制了对事实性与一致性的全面定义。本文提出 TeCFaP(Temporal Consistent Factuality Probe)任务,拓展时序维度的一致性评测。为此,构建高质量的 TEMP-COFAC 数据集,包含前缀式英语查询改写。同时扩展现有指标以刻画跨时间维度的一致性事实性。实验表明,多数大模型在 TeCFaP 上表现较差。为此,提出 CoTSeLF(Consistent-Time-Sensitive Learning Framework),结合多任务指令微调(MT-IT)与一致时间敏感强化学习(CTSRL),有效提升模型在时间一致性上的表现。实验证明其优于多个基线。

原文摘要 · Abstract (English)

The prolific use of Large Language Models (LLMs) as an alternate knowledge base requires them to be factually consistent, necessitating both correctness and consistency traits for paraphrased queries. Recently, significant attempts have been made to benchmark datasets and metrics to evaluate LLMs for these traits. However, structural simplicity (subject-relation-object) and contemporary association in their query formulation limit the broader definition of factuality and consistency. In this study, we introduce TeCFaP, a novel Temporally Consistent Factuality Probe task to expand the consistent factuality probe in the temporal dimension. To this end, we propose TEMP-COFAC, a high-quality dataset of prefix-style English query paraphrases. Subsequently, we extend the definitions of existing metrics to represent consistent factuality across temporal dimension. We experiment with a diverse set of LLMs and find most of them performing poorly on TeCFaP. Next, we propose a novel solution CoTSeLF (Consistent-Time-Sensitive Learning Framework) combining multi-task instruction tuning (MT-IT) with consistent-time-sensitive reinforcement learning (CTSRL) to improve temporally consistent factuality in LLMs. Our experiments demonstrate the efficacy of CoTSeLF over several baselines.

大模型评测事实性时序一致性训练框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。