通过反转概念检测大模型推理鲁棒性,提出新方法提升逻辑一致性。
Concept-Reversed Winograd Schema Challenge: Evaluating and Improving Robust Reasoning in Large Language Models via Abstraction
- 反转语义关联构造对抗性题目,暴露模型表面推理缺陷。
- 大模型在新数据集上性能下降超40%,验证其依赖表层线索。
- 提出抽象思维提示法,显著恢复模型推理稳定性,适合可信AI研究者。
尽管大型语言模型在推理方面表现出色,但其仍存在因语义关联和浅层逻辑链导致的幻觉与不可靠推理问题。为评估大模型是否具备稳健推理能力,我们基于著名的温格罗德模式挑战(Winograd Schema Challenge, WSC)数据集,提出了一个新的评估数据集——概念反转温格罗德模式挑战(CR-WSC)。通过简单地将原题中与错误答案更相关的概念进行反转,我们发现尽管推理逻辑保持不变,大模型性能却显著下降。此外,我们提出了新的提示方法——思维抽象(Abstraction-of-Thought, AoT),利用概念抽象将对抗性案例还原为正常案例,实验表明该方法能有效提升大模型在CR-WSC上的推理鲁棒性与一致性。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have showcased remarkable proficiency in reasoning, there is still a concern about hallucinations and unreliable reasoning issues due to semantic associations and superficial logical chains. To evaluate the extent to which LLMs perform robust reasoning instead of relying on superficial logical chains, we propose a new evaluation dataset, the Concept-Reversed Winograd Schema Challenge (CR-WSC), based on the famous Winograd Schema Challenge (WSC) dataset. By simply reversing the concepts to those that are more associated with the wrong answer, we find that the performance of LLMs drops significantly despite the rationale of reasoning remaining the same. Furthermore, we propose Abstraction-of-Thought (AoT), a novel prompt method for recovering adversarial cases to normal cases using conceptual abstraction to improve LLMs' robustness and consistency in reasoning, as demonstrated by experiments on CR-WSC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。