用数学操演理论构建大模型推理的严谨框架,提升多步问答可靠性。
Operads for compositional reasoning in LLMs
- 用操演结构建模问题分解与答案组合,提供数学基础。
- 提出操演一致性指标,与模型准确率强相关,优于传统方法。
- 适合研究大模型推理机制与提升可靠性的学者参考。
问题分解——将复杂问题拆分为更简单的子问题并组合其答案以获得最终答案——是提升大模型推理能力的常用策略,但目前缺乏严格的数学基础。本文提出操演(operads)作为描述问题分解的自然框架,定义了问题操演 $Q$,其中操作对应问题模板,组合对应子答案代入。我们证明问答模型可被解释为 $Q$ 上的代数。这一视角不仅重构现有实践,还引出新方法,特别是操演一致性概念,用于衡量模型在问题分解树部分坍缩下的答案一致性。在配套论文(Bottman, Liu, and Richardson, 2026)中,对十二个大模型和四个多跳问答数据集的实证评估表明,操演一致性与准确率强相关,并优于标准温度自洽基线。我们主张操演是问题分解的自然数学家园,其不变量如操演一致性为分析和改进多步推理的可靠性开辟了新方向。
原文摘要 · Abstract (English)
Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM reasoning, yet it currently lacks a rigorous mathematical foundation. In this paper, we propose operads, mathematical structures that model many-in, one-out operations and compositions thereof, as a natural framework for describing question decomposition. We define the questions operad $Q$, in which operations correspond to question templates and composition corresponds to substitution of sub-answers, and show how QA models can be interpreted as algebras over $Q$. Beyond reframing existing practice, this operadic perspective points toward new methods, in particular a notion of operadic consistency, which measures whether a QA model's answers agree across the partial collapses of a question decomposition tree. Empirical evaluation of operadic consistency is reported in our companion paper (Bottman, Liu, and Richardson, 2026), which finds it strongly correlated with accuracy across twelve LLMs and four multi-hop QA datasets and outperforming standard temperature-based self-consistency baselines. We argue that operads are the natural mathematical home for question decomposition, and that invariants such as operadic consistency open new directions for analyzing and improving the reliability of multi-step reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。