首个日语信念不一致逻辑推理数据集,检验大模型是否受直觉误导。
BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic Reasoning
- 构建日语逻辑题库,专测模型在合理但反直觉结论下的判断力
- 顶尖模型准确率超99%,而部分日语专用模型不足60%
- 提示词设计和推理深度显著影响模型抗偏见能力
我们提出 BIS Reasoning 1.0,首个专为评估大语言模型(LLMs)在信念不一致情境下逻辑推理能力而设计的大规模日语数据集。与 NeuBAROCO、JFLD 等强调通用或信念一致逻辑的资源不同,BIS Reasoning 1.0 系统性引入逻辑有效但违背常识的三段论,以揭示信念偏差——即模型倾向于接受看似可信但未必正确的结论。我们在统一零样本协议下测试包括 OpenAI GPT-5 变体、GPT-4o、Qwen 及主流日语 LLM 在内的代表性模型。以推理为核心优化的模型表现优异,如 Qwen3-32B 达约 99%,GPT-5-mini 高达约 99.7%;而 GPT-4o 约 80%;早期日语专用模型普遍低于 60%,最新版 llm-jp-3.1-13b-instruct4 提升至中等 80% 水平。结果表明,抵抗信念冲突的能力主要依赖显式推理优化,而非语言专精或规模。分析显示,即使顶级系统在逻辑与直觉/事实冲突时仍会出错,且性能对提示设计与推理投入敏感。该研究对法律、医疗、科学文献等需严格逻辑可靠性的领域具有重要启示。
原文摘要 · Abstract (English)
We present BIS Reasoning 1.0, the first large-scale Japanese dataset of syllogistic reasoning problems explicitly designed to evaluate belief-inconsistent reasoning in large language models (LLMs). Unlike prior resources such as NeuBAROCO and JFLD, which emphasize general or belief-aligned logic, BIS Reasoning 1.0 systematically introduces logically valid yet belief-inconsistent syllogisms to expose belief bias, the tendency to accept believable conclusions irrespective of validity. We benchmark a representative suite of cutting-edge models, including OpenAI GPT-5 variants, GPT-4o, Qwen, and prominent Japanese LLMs, under a uniform, zero-shot protocol. Reasoning-centric models achieve near-perfect accuracy on BIS Reasoning 1.0 (e.g., Qwen3-32B $\approx$99% and GPT-5-mini up to $\approx$99.7%), while GPT-4o attains around 80%. Earlier Japanese-specialized models underperform, often well below 60%, whereas the latest llm-jp-3.1-13b-instruct4 markedly improves to the mid-80% range. These results indicate that robustness to belief-inconsistent inputs is driven more by explicit reasoning optimization than by language specialization or scale alone. Our analysis further shows that even top-tier systems falter when logical validity conflicts with intuitive or factual beliefs, and that performance is sensitive to prompt design and inference-time reasoning effort. We discuss implications for safety-critical domains, including law, healthcare, and scientific literature, where strict logical fidelity must override intuitive belief to ensure reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。