构建逻辑组合式常识推理基准,揭示模型在否定推理上的显著短板。
LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
- 用与/或/非等逻辑操作重构常识推理任务
- 模型对否定类问题准确率大幅下降,暴露核心缺陷
- 适合研究常识推理、模型可解释性与逻辑能力的学者
常识推理常需评估多个合理解释而非单一答案,但现有基准多采用单标签评价,难以区分陈述是否共存合理、互斥或均不合理。我们提出LOGICAL-COMMONSENSEQA,将常识推理重新定义为基于原子陈述的逻辑组合,使用可解释的可判定性算子(AND、OR、NEITHER/NOR)。在零样本、少样本及思维链提示下评估指令微调、推理专用和微调模型,发现模型在合取和部分析取推理上表现尚可,但在否定类问题上性能急剧下降。该基准揭示了模型在复合常识推理中的根本局限,并提供了一个可控框架以推动该领域发展。
原文摘要 · Abstract (English)
Commonsense reasoning often involves evaluating multiple plausible interpretations rather than selecting a single atomic answer, yet most benchmarks rely on single-label evaluation, obscuring whether statements are jointly plausible, mutually exclusive, or jointly implausible. We introduce LOGICAL-COMMONSENSEQA, a benchmark that reframes commonsense reasoning as logical composition over pairs of atomic statements using plausibility-level operators (AND, OR and NEITHER/NOR). Evaluating instruction-tuned, reasoning-specialized, and fine-tuned models under zero-shot, few-shot, and chain-of-thought prompting, we find that while models perform reasonably on conjunctive and moderately on disjunctive reasoning, performance degrades sharply on negation-based questions. LOGICAL-COMMONSENSEQA exposes fundamental reasoning limitations and provides a controlled framework for advancing compositional commonsense reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。