arXiv:2601.01982cs.AIcs.IT2026-01中稿 · AAAI被引 1

测试大模型在混沌系统中的逻辑推理能力,发现其虽准确但缺乏连贯性。

ChaosBench-Logic: A Benchmark for Logical and Symbolic Reasoning on Chaotic Dynamical Systems

  • 用一阶逻辑统一建模30个混沌系统,标注11类语义谓词
  • 大模型单题准确率达91%-94%,但组合推理全错、对话一致性差
  • 适合研究科学推理、神经符号系统或模型可解释性的学者

大型语言模型在自然语言任务中表现优异,但在需要精确逻辑与符号推理的领域仍显脆弱。混沌动力系统因其确定性却常被误认为随机性或复杂性,构成严峻挑战。我们提出ChaosBench-Logic基准,基于统一的一阶逻辑(FOL)本体,评估大模型在30个多样化动力系统上的推理能力。每个系统标注11个语义谓词的真值,生成621道涵盖多跳推理、跨系统类比、反事实推理、偏见探测及多轮对话等七类问题。定义逻辑准确性、蕴含一致性、对话连贯性与矛盾率等指标,并开源评估流水线。初步实验显示,前沿模型如GPT-4、Claude 3.5 Sonnet、Gemini 2.5 Flash和LLaMA-3 70B在单题上达到91%-94%准确率,但在组合式任务中得分为0%,全局连贯性脆弱;对话级准确率介于53.1%(GPT-4 CoT)至75.5%(LLaMA-3零样本)。ChaosBench-Logic为诊断模型缺陷提供严谨测试平台,也为发展提升科学推理的神经符号方法奠定基础。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at natural language tasks but remain brittle in domains requiring precise logical and symbolic reasoning. Chaotic dynamical systems provide an especially demanding test because chaos is deterministic yet often misinterpreted as randomness or complexity. We introduce ChaosBench-Logic, a benchmark that evaluates LLM reasoning across 30 diverse dynamical systems using a unified first-order logic (FOL) ontology. Each system is annotated with truth assignments for 11 semantic predicates, and 621 questions are generated across seven reasoning categories, including multi-hop implications, cross-system analogies, counterfactual reasoning, bias probes, and multi-turn dialogues. We define metrics for logical accuracy, implication consistency, dialogue coherence, and contradiction, and we release an open-source evaluation pipeline. Initial experiments show that frontier LLMs such as GPT-4, Claude 3.5 Sonnet, Gemini 2.5 Flash, and the open-source LLaMA-3 70B achieve 91-94% per-item accuracy, yet still score 0% on compositional items and exhibit fragile global coherence. Dialogue-level accuracy ranges from 53.1% (GPT-4 CoT) to 75.5% (LLaMA-3 zero-shot). ChaosBench-Logic provides a rigorous testbed for diagnosing such failures and a foundation for developing neuro-symbolic approaches that improve scientific reasoning in LLMs.

逻辑推理混沌系统大模型评测神经符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。