arXiv:2604.23072cs.AI2026-04被引 2

用软命题推理提升大模型分析的稳定性和可扩展性

Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis

论文配图:Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis
图 1 · 摘自论文原文
  • 将复杂分析拆解为命题评估,通过分治框架降低误差
  • 在金融政治预测中平均提升15.84%准确率,方差仅6.02%
  • 适合需要高可靠性与可交互分析的科研和决策场景

大型语言模型(LLM)代理在金融预测、科学发现等复杂分析任务中日益重要,但其推理存在随机不稳定性且缺乏可验证的组合结构。为此,我们提出Analytica,一种基于软命题推理(SPR)的新架构。SPR将复杂分析重构为对不同结果命题的软真值估计过程,可形式化建模并最小化估计误差的偏差与方差。Analytica采用并行分治框架:为降偏差,将问题分解为子命题树,并引入工具增强的LLM奠基代理,包括一个用于数据驱动分析的新型Jupyter Notebook代理来验证和打分事实;为降方差,递归使用稳健线性模型合成这些已奠基的叶节点,有效平均随机噪声,实现更高效率、可扩展性及交互式“若-则”分析能力。理论与实证结果表明,在经济、金融和政治预测任务中,Analytica相较多样基础模型平均提升15.84%准确率,使用Deep Research奠基代理时达到71.06%准确率,方差最低仅6.02%。其Jupyter Notebook奠基代理表现出强成本效益,以90.35%更低成本、52.85%更短时间实现70.11%准确率。Analytica在分析深度增加时仍具高度抗噪性与稳定性能增长,近似线性时间复杂度,并良好适配开源权重模型与科学领域。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet their reasoning suffers from stochastic instability and lacks a verifiable, compositional structure. To address this, we introduce Analytica, a novel agent architecture built on the principle of Soft Propositional Reasoning (SPR). SPR reframes complex analysis as a structured process of estimating the soft truth values of different outcome propositions, allowing us to formally model and minimize the estimation error in terms of its bias and variance. Analytica operationalizes this through a parallel, divide-and-conquer framework that systematically reduces both sources of error. To reduce bias, problems are first decomposed into a tree of subpropositions, and tool-equipped LLM grounder agents are employed, including a novel Jupyter Notebook agent for data-driven analysis, that help to validate and score facts. To reduce variance, Analytica recursively synthesizes these grounded leaves using robust linear models that average out stochastic noise with superior efficiency, scalability, and enable interactive "what-if" scenario analysis. Our theoretical and empirical results on economic, financial, and political forecasting tasks show that Analytica improves 15.84% accuracy on average over diverse base models, achieving 71.06% accuracy with the lowest variance of 6.02% when working with a Deep Research grounder. Our Jupyter Notebook grounder shows strong cost-effectiveness that achieves a close 70.11% accuracy with 90.35% less cost and 52.85% less time. Analytica also exhibits highly noise-resilient and stable performance growth as the analysis depth increases, with a near-linear time complexity, as well as good adaptivity to open-weight LLMs and scientific domains.

大模型推理命题推理分析系统可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。