用关系复杂度测试大模型,发现高阶关系推理能力普遍不足
Evaluating Relational Reasoning in LLMs with REL

- 引入关系复杂度(RC)衡量多实体绑定的难度
- 模型性能随RC上升而持续下降,即使实体总数不变
- 适合关注模型深层推理能力的研究者参考
关系推理是推断多个实体、属性或变量之间关联的能力,对科学推理至关重要。现有评估多集中于表格、图结构或合成任务,未分离高阶关系绑定带来的难度。本文提出关系复杂度(RC)——需同时绑定的独立实体或操作数最小数量,作为控制输入规模、词汇和表征等干扰因素的原理性指标。基于此,构建涵盖代数、化学、生物学的生成式基准框架REL,各领域内变化RC。在前沿大模型上,性能随RC增加而单调下降,即使实体总数固定。该缺陷在增加推理计算量和上下文学习后仍存在,表明问题根源在于关系绑定的阶数,而非推理步数不足或缺乏示例。研究揭示当前模型在高阶关系推理中的困境,呼吁以关系复杂度重新审视评估基准。
原文摘要 · Abstract (English)
Relational reasoning is the ability to infer relations that jointly bind multiple entities, attributes, or variables. This ability is central to scientific reasoning, but existing evaluations of relational reasoning in large language models often focus on structured inputs such as tables, graphs, or synthetic tasks, and do not isolate the difficulty introduced by higher-arity relational binding. We study this problem through the lens of Relational Complexity (RC), which we define as the minimum number of independent entities or operands that must be simultaneously bound to apply a relation. RC provides a principled way to vary reasoning difficulty while controlling for confounders such as input size, vocabulary, and representational choices. Building on RC, we introduce REL, a generative benchmark framework spanning algebra, chemistry, and biology that varies RC within each domain. Across frontier LLMs, performance degrades consistently and monotonically as RC increases, even when the total number of entities is held fixed. This failure mode persists with increased test-time compute and in-context learning, suggesting a limitation tied to the arity of the required relational binding rather than to insufficient inference steps or lack of exposure to examples. Our results identify a regime of higher-arity reasoning in which current models struggle, and motivate re-examining benchmarks through the lens of relational complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。