测试大模型在跨学科科学推理中的能力边界,发现组合顺序越复杂越容易崩溃。
XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition

- 设计跨领域交互式推理任务,模拟真实科研流程。
- 8598个会话显示组合顺序增加导致推理失败率显著上升。
- 适合研究AI在复杂科学任务中可靠性与错误传播的学者。
大型语言模型(LLMs)被越来越多地用于知识综合,但其在科学知识组合泛化方面的能力仍缺乏系统评估。现有基准多聚焦于单轮受限场景,无法反映真实交互式科研工作流中的能力极限。为此,我们提出XDomainBench,一个用于交互式跨学科科学推理的诊断性基准。通过形式化组合顺序与混合结构,实现从单一学科到跨学科的系统性压力测试,涵盖20个领域、4类任务和8种真实轨迹模式,共8598个交互会话,模拟真实AI for Science(AI4S)场景。大规模评估发现,随着组合顺序复杂度增加,模型出现系统性推理崩溃,根源在于两点:(i) 领域组合直接导致难度上升;(ii) 轨迹模式引发错误累积、推理中断与领域混淆,最终造成会话崩溃。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundaries exposed by real-world interactive scientific workflows. To address this, we introduce XDomainBench, a diagnostic benchmark for interactive interdisciplinary scientific reasoning. We formalize the composition order and mixture structure to enable systematic stress-testing from single-discipline to inter-disciplinary, comprising 8,598 interactive sessions across 20 domains and 4 task categories, with 8 realistic trajectory patterns covering difficulty and domain-mixture dynamics, simulating real AI4S scenarios. Large-scale evaluation of LLMs reveals a systematic reasoning collapse as composition order increases, stemming from two root causes: (i) direct difficulty increases induced by domain composition, and (ii) indirect interaction-amplified failures where trajectory patterns trigger error accumulation, reasoning breaks, and domain confusion, ultimately leading to session collapse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。