提出新评测集CORE,检验大模型区分语义关系与无关性的能力。
CORE: Comprehensive Ontological Relation Evaluation for Large Language Models
- 构建225K题目的跨学科多选题数据集,涵盖74个领域
- 29个顶尖大模型在无关对上准确率仅0-41.35%,平均下降超30%
- 揭示模型对无关关系的误判问题,适合评估模型安全性和推理可靠性
大型语言模型在诸多推理基准上表现良好,但现有评测极少考察其区分有意义语义关系与真正无关性能力。我们提出CORE(综合本体关系评测),包含225,000道覆盖74个学科的多选题,并提供一个203道严格验证的通用领域开源基准(Cohen's Kappa = 1.0),涵盖24种语义关系类型,无关对均衡分布。人类基线来自1000+参与者,整体准确率达92.6%(无关对95.1%)。相比之下,29个最先进的大模型整体准确率为48.25%-70.9%,相关对表现接近完美(86.5%-100%),但在无关对上严重退化至0-41.35%,尽管置信度仍高达92%-94%。无关对上的期望校准误差增加2-4倍,平均语义坍缩率达37.6%,表明系统性生成虚假关系。在CORE 225K多选题数据集上,准确率进一步降至约2%,凸显领域特定语义推理的巨大挑战。我们识别出无关性推理是大模型评测与安全的关键未被充分评估前沿。
原文摘要 · Abstract (English)
Large Language Models (LLMs) perform well on many reasoning benchmarks, yet existing evaluations rarely assess their ability to distinguish between meaningful semantic relations and genuine unrelatedness. We introduce CORE (Comprehensive Ontological Relation Evaluation), a dataset of 225K multiple-choice questions spanning 74 disciplines, together with a general-domain open-source benchmark of 203 rigorously validated questions (Cohen's Kappa = 1.0) covering 24 semantic relation types with equal representation of unrelated pairs. A human baseline from 1,000+ participants achieves 92.6% accuracy (95.1% on unrelated pairs). In contrast, 29 state-of-the-art LLMs achieve 48.25-70.9% overall accuracy, with near-ceiling performance on related pairs (86.5-100%) but severe degradation on unrelated pairs (0-41.35%), despite assigning similar confidence (92-94%). Expected Calibration Error increases 2-4x on unrelated pairs, and a mean semantic collapse rate of 37.6% indicates systematic generation of spurious relations. On the CORE 225K MCQs dataset, accuracy further drops to approximately 2%, highlighting substantial challenges in domain-specific semantic reasoning. We identify unrelatedness reasoning as a critical, under-evaluated frontier for LLM evaluation and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。