测试大模型在多线索并行推理中的综合能力,发现深度推理不等于广度推理。
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

- 设计新基准MPAR-Bench,通过多线索协同推理解密目标
- 模型在扰动下准确率下降5-18个百分点,显示鲁棒性不足
- 适合研究推理泛化与模型可解释性的研究人员参考
大型语言模型在需要长而复杂推理链的任务上取得了显著进展,这主要体现为推理深度的提升。一种互补但未被充分探索的能力是推理广度:并行探索多个语义方向,并将所得线索整合为一致答案。我们提出了MPAR-Bench,一个中英文双语基准,通过多点关联推理来分离推理广度。受合作游戏《Just One》启发,每道题要求模型从若干独立生成、语义多样化的线索中恢复隐藏目标。我们使用多智能体线索生成管道、基于嵌入的多样性过滤和人工验证构建了1,000个题目,仅答案空间来自公开词表,所有线索集均从零生成。除精确匹配准确率外,还评估了准确率、ANLS、嵌入相似性、推理轨迹验证及四种扰动:线索遮蔽、顺序打乱、干扰项注入和多步线索。在不同模型中,扰动使英语准确率下降9-18个百分点,中文下降5-12个百分点。思维模式提升了标准表现,尤其在英语中,但并未一致降低对扰动的敏感性。案例分析显示,延长推理可能推翻初始正确假设。结果表明,更强的推理深度并未自动带来鲁棒的推理广度,且当前基准尚未充分覆盖这一能力。
原文摘要 · Abstract (English)
Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reasoning depth. A complementary and comparatively unexamined capability is reasoning breadth: exploring multiple semantic directions in parallel and integrating the resulting clues into one coherent answer. We introduce MPAR-Bench, a bilingual English-Chinese benchmark that isolates reasoning breadth through multi-point associative reasoning. Inspired by the cooperative game Just One, each item asks a model to recover a hidden target from several independently generated, semantically diverse clues. We construct 1,000 items using a multi-agent clue-generation pipeline, embedding-based diversity filtering, and human verification. Only the answer space is drawn from public word lists, whereas every clue set is generated from scratch. Beyond exact-match accuracy, we evaluate models using accuracy, ANLS, embedding similarity, reasoning-trace verification, and four perturbations: clue masking, order shuffling, distractor injection, and multi-step clues. Across evaluated models, perturbations reduce accuracy by 9-18 percentage points in English and 5-12 percentage points in Chinese. Thinking mode improves standard-setting accuracy, especially in English, but does not consistently reduce sensitivity to perturbations. Case-level analysis also shows that extended reasoning can overturn an initially correct hypothesis. These results indicate that greater reasoning depth does not automatically confer robust reasoning breadth, and that reasoning breadth remains largely uncovered by current benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。