用结构化指令图提升大模型推理效率与成本表现
BRAID: Bounded Reasoning for Autonomous Inference and Decisions
- 采用可机器解析的Mermaid指令图实现约束式推理
- 在多个数据集上准确率提升且推理成本显著降低
- 适合需要高效决策的生产级智能体系统
大型语言模型在性能、成本与令牌消耗之间呈现非线性关系。本文通过BRAID(受限推理自主推断与决策)框架,对多层级GPT模型进行了结构化提示的定量研究,评估涵盖AdvancedIF、GSM-Hard及SCALE MultiChallenge基准数据集。BRAID引入基于Mermaid的指令图结构化推理框架,使模型以结构性方式推理,而非依赖无界自然语言扩展。结果表明,结构化可机器读取的提示能显著提升智能体在生产系统中的推理准确率与成本效率。研究证实BRAID是优化自主代理系统推理效率的有效且可扩展的方法。所有数据集与详细结果日志已公开于https://benchmark.openserv.ai。
原文摘要 · Abstract (English)
Large Language Models (LLMs) exhibit nonlinear relationships between performance, cost, and token usage. This paper presents a quantitative study on structured prompting using BRAID (Bounded Reasoning for Au tonomous Inference and Decisions) across multiple GPT model tiers, eval uated on the AdvancedIF, GSM-Hard, and the SCALE MultiChallenge benchmark datasets. BRAID introduces a bounded reasoning framework using Mermaid-based instruction graphs that enable models to reason struc turally rather than through unbounded natural-language token expansion. We show that structured machine-readable prompts substantially increase reasoning accuracy and cost efficiency for agents in production systems. The findings establish BRAID as an effective and scalable technique for optimizing inference efficiency in autonomous agent systems. All datasets and detailed result logs are available at https://benchmark.openserv.ai.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。