构建可验证的逻辑推理数据集,让大模型逻辑能力经得起严格检验。
FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale
- 用程序自动验证生成大规模逻辑推理轨迹
- 模型在关键任务上准确率仅45.7%和27%
- 适合评估大模型逻辑推理真实水平
语言模型的推理难以评估:自然语言推理链不可验证,符号化数据集规模太小,多数基准混淆启发式与真正推理。我们提出FOL-Traces,首个大规模、程序化验证的逻辑推理轨迹数据集,支持对结构化逻辑推理进行严谨评估。我们设计两项挑战性且全面的诊断任务——掩码操作预测与步骤补全,直接探测模型的语法意识与过程忠实度。在5个推理型大模型上系统实验表明,该数据集仍具挑战性:模型在掩码操作预测任务上准确率约为45.7%,在双步补全任务上约为27%。
原文摘要 · Abstract (English)
Reasoning in language models is difficult to evaluate: natural-language traces are unverifiable, symbolic datasets are too small, and most benchmarks conflate heuristics with inference. We present FOL-Traces, the first large-scale dataset of programmatically verified reasoning traces, enabling rigorous evaluation of structured logical inference. We also propose two challenging and comprehensive diagnostic tasks-masked operation prediction and step completion-that directly probe syntactic awareness and process fidelity. FOL-Traces serves as a scalable testbed for rigorously studying how models perform structured logical inference. Systematic experiments with 5 reasoning LLMs show that the dataset remains challenging: models only reach around 45.7% accuracy on masked operation prediction and around 27% on two-step completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。