构建首个评估大模型抗逻辑谬误能力的基准,揭示不同模型的脆弱性差异。
Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies

- 通过多智能体生成事实问题与谬论配对,构建真实对抗场景。
- 发现大模型在不同谬论类型下表现差异显著,抗性不一。
- 提出新指标LFR@k,分离知识局限与谬论抗性,适合安全评测研究者。
大型语言模型(LLMs)具备强大的语义能力,但其对逻辑谬误等操纵性语言模式的鲁棒性尚未深入研究。以往工作多关注模型能否识别或分类谬误,而对其在持续恶意说服下的稳健性探讨不足。为此,我们提出LoFa(Logical Fallacy),一个全面评估大模型抗谬论能力的基准。LoFa基于多智能体流水线,将事实性问题与谬论论证配对,并引入多轮辩论框架,以评估模型在持续对抗性说服下的表现。为区分谬论抗性与模型固有知识限制,我们进一步提出逻辑谬论抵抗度指标LFR@k,量化模型抵御谬论攻击的能力。实验表明,不同大模型在各类谬论中表现出不同的鲁棒性水平,揭示了各模型独特的脆弱性特征。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies remains underexplored. Prior work has primarily examined whether LLMs can identify or classify fallacies, leaving their robustness against fallacious persuasion insufficiently studied. To address this gap, we introduce LoFa (Logical Fallacy), a comprehensive benchmark for evaluating LLM robustness against fallacies. LoFa is constructed through a multi-agent pipeline that pairs factual questions with fallacious arguments, and is accompanied by a multi-round debate framework for assessing model resilience under sustained adversarial persuasion. To disentangle fallacy robustness from a model's inherent knowledge limitations, we further propose Logical Fallacy Resistance at k (LFR@k), a metric that quantifies resistance to fallacious attacks. Experiments show that LLMs exhibit varying levels of robustness across different fallacy types, revealing distinct vulnerability profiles among models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。