用追问检验大模型解释的一致性,发现其推理漏洞。
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
- 结合符号提取与语言模型生成追问问题。
- 比纯大模型生成的追问更准确、多样。
- 适合研究模型可解释性与可信AI的学者。
大型语言模型(LLMs)常被要求解释其输出以提升准确性和透明度。然而,研究表明这些解释可能歪曲模型的真实推理过程。一种有效识别解释中错误或遗漏的方法是通过一致性检查,通常涉及提出后续问题。本文提出Cross-Examiner,一种基于模型对初始问题解释生成后续问题的新方法。该方法结合符号信息提取与语言模型驱动的问题生成,生成的追问问题优于仅依赖大模型生成的结果。此外,该方法更具灵活性,能生成更广泛类型的追问问题。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are often asked to explain their outputs to enhance accuracy and transparency. However, evidence suggests that these explanations can misrepresent the models' true reasoning processes. One effective way to identify inaccuracies or omissions in these explanations is through consistency checking, which typically involves asking follow-up questions. This paper introduces, cross-examiner, a new method for generating follow-up questions based on a model's explanation of an initial question. Our method combines symbolic information extraction with language model-driven question generation, resulting in better follow-up questions than those produced by LLMs alone. Additionally, this approach is more flexible than other methods and can generate a wider variety of follow-up questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。