arXiv:2602.12889cs.CL2026-02

评测大模型在命理符号与时间组合推理中的表现,揭示其短板。

BaziQA-Benchmark: Evaluating Symbolic and Temporally Compositional Reasoning in Large Language Models

  • 基于命理竞赛题构建标准化评测集,需结合符号图表与时间条件
  • 模型表现远超随机但对时间组合敏感,精准定位与多条件判断仍弱
  • 引入轻量推理协议,可控制推理顺序以分析行为机制

我们提出 BaziQA-Benchmark,一个用于评估大语言模型在符号与时间复合推理能力上的标准化基准。该基准源自全球命理师大赛(2021–2025)中200道专业筛选的多选题,每题需在固定符号图谱和相互作用的时间条件下进行结构化推理。与非结构化或提示驱动的评估不同,BaziQA-Benchmark支持客观评分,并可在年份、领域和模型家族间实现可控对比。我们在多轮对话设置下评估主流语言模型,分析其在时间难度、推理领域及推理协议下的性能差异。为进一步探究推理行为,我们引入一种轻量级结构化推理协议,约束推理顺序但不增加领域知识。结果表明,模型虽持续优于随机猜测,但仍远未达到饱和,对时间组合与推理顺序高度敏感,且在精确时间定位和多条件符号判断上存在系统性失败。

原文摘要 · Abstract (English)

We present BaziQA-Benchmark, a standardized benchmark for evaluating symbolic and temporally compositional reasoning in large language models. The benchmark is derived from 200 professionally curated, multiple-choice problems from the Global Fortune-teller Competition (2021--2025), where each instance requires structured inference over a fixed symbolic chart and interacting temporal conditions. Unlike anecdotal or prompt-driven evaluations, BaziQA-Benchmark enables objective scoring and controlled comparison across years, domains, and model families. We evaluate contemporary language models under a multi-turn setting and analyze performance variation across temporal difficulty, reasoning domains, and inference protocols.To further probe reasoning behavior, we introduce a lightweight Structured Reasoning Protocol that constrains inference order without adding domain knowledge. Results show that models consistently outperform chance but remain far from saturation, exhibiting pronounced sensitivity to temporal composition and reasoning order, as well as systematic failures on precise temporal localization and multi-condition symbolic judgments.

推理评估符号推理时间推理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。