arXiv:2505.21115cs.CL2025-05EMNLP被引 6

区分问题是否随时间变化,提升大模型问答可信度

Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA

  • 构建首个多语言永续问题数据集EverGreenQA
  • 发现大模型多数依赖不确定性信号而非显式判断
  • 可改进知识评估、数据过滤和模型解释

大语言模型在问答任务中常产生幻觉,一个重要但未被充分研究的因素是问题的时间属性——即答案是否随时间保持稳定(永续)或会变化(可变)。本文提出EverGreenQA,首个带有永续标签的多语言问答数据集,支持评估与训练。基于该数据集,我们对12个主流大模型进行评测,分析其是否通过显式语义判断或隐式不确定信号来捕捉问题时间性。同时,我们训练了EG-E5,一个轻量级多语言分类器,在该任务上达到当前最优性能。最后,我们展示了永续分类在三个实际场景中的应用价值:改进自我知识估计、过滤问答数据集、解释GPT-4o检索行为。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often hallucinate in question answering (QA) tasks. A key yet underexplored factor contributing to this is the temporality of questions -- whether they are evergreen (answers remain stable over time) or mutable (answers change). In this work, we introduce EverGreenQA, the first multilingual QA dataset with evergreen labels, supporting both evaluation and training. Using EverGreenQA, we benchmark 12 modern LLMs to assess whether they encode question temporality explicitly (via verbalized judgments) or implicitly (via uncertainty signals). We also train EG-E5, a lightweight multilingual classifier that achieves SoTA performance on this task. Finally, we demonstrate the practical utility of evergreen classification across three applications: improving self-knowledge estimation, filtering QA datasets, and explaining GPT-4o retrieval behavior.

问答系统多语言可信AI时间性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。