模型事实知识稳定,但组合推理能力仍可能很差,现有评估方法会掩盖这一问题。
Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning

- 设计双门控协议,分离原子知识与组合推理能力
- 同一数据集上组合表现相差超40个百分点,源于评估方式缺陷
- 适合关注多跳推理真实能力的研究者与评测人员
后训练通常通过综合基准分数评估,将多跳推理视为单一能力——认为答对更多题的模型必然更擅长整合事实。我们发现这种假设具有误导性:在统计上原子知识相近的四种后训练方案,其组合行为差异超过40个百分点,这一现象称为组合坍缩:即稳定掌握的事实无法被有效串联成推理链,而这一问题在综合指标下完全不可见。本文提出双门控协议,将评估目标从整体组合差距转变为在原子知识稳定条件下的残余组合失败,将后训练提升分解为三个独立通道:原子稳定性、残余组合能力与关键深度。在涵盖深度2至11的四个时间事实链基准上,该分解揭示了后训练目标实际改变了组合能力方向,而这些变化被综合指标所掩盖。研究建议,关于多跳推理改进的结论应辅以原子门控的组合度量。诊断探针进一步表明,大量组合失败源于生成时计算约束,而非永久性的组合能力缺失。
原文摘要 · Abstract (English)
Post-training is routinely evaluated through aggregate benchmark scores that treat multi-hop reasoning as a single capability -- as if a model that answers more questions correctly must be better at assembling facts. We show that this assumption can be misleading: recipes with statistically indistinguishable atomic knowledge produce composition behaviour separated by over 40 percentage points, a phenomenon we call composition collapse: the systematic failure to assemble stably-known facts into chains, invisible to aggregate metrics. We introduce a double-gate protocol that changes the estimand from an aggregate compositionality gap to residual composition failure conditioned on stable atomic access, decomposing post-training gains into three independent channels: atomic stability, residual composition, and critical depth. On a benchmark of temporal factual chains spanning depths 2--11 across four post-training recipes, this decomposition reveals that post-training objectives shift composition capability in directions that aggregate metrics mask, and suggests that claims about multi-hop reasoning improvement should be accompanied by atomic-gate-controlled composition metrics. Diagnostic probes further show that a substantial share of measured composition failure reflects generation-time computation constraints rather than permanent inability to compose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。