用可追溯的LLM流程模拟专家决策,让AI推理过程透明可查。
LLM Driven Processes to Foster Explainable AI
- 分模块设计,整合敏感性分析、博弈论等三种决策框架。
- 100次物流测试中,因子对齐率达55.5%,角色匹配率57%。
- 适合需要可解释性的人工智能应用,如医疗、金融决策支持。
我们提出一种模块化、可解释的LLM代理决策支持系统,将推理过程转化为可审计的中间产物。系统实现三种框架:维斯特的敏感性模型(因子集、带符号影响矩阵、系统角色、反馈回路)、正则形式博弈(策略、收益矩阵、均衡点)以及序贯博弈(角色条件代理、树结构构建、逆向归纳)。每一步均可替换模块。使用默认的GPT-5 LLM配合确定性分析器进行均衡计算与基于矩阵的角色分类,生成可追踪的中间结果而非黑箱输出。在真实物流案例中(100次运行),26个因子的平均因子对齐率为55.5%,运输核心子集达62.9%;角色匹配一致率为57%。采用八项标准评分卡(最高100分)的LLM评判者得分与重构的人类基准相当。因此,可配置的LLM流水线能以透明、可检查的步骤模拟专家工作流。
原文摘要 · Abstract (English)
We present a modular, explainable LLM-agent pipeline for decision support that externalizes reasoning into auditable artifacts. The system instantiates three frameworks: Vester's Sensitivity Model (factor set, signed impact matrix, systemic roles, feedback loops); normal-form games (strategies, payoff matrix, equilibria); and sequential games (role-conditioned agents, tree construction, backward induction), with swappable modules at every step. LLM components (default: GPT-5) are paired with deterministic analyzers for equilibria and matrix-based role classification, yielding traceable intermediates rather than opaque outputs. In a real-world logistics case (100 runs), mean factor alignment with a human baseline was 55.5\% over 26 factors and 62.9\% on the transport-core subset; role agreement over matches was 57\%. An LLM judge using an eight-criterion rubric (max 100) scored runs on par with a reconstructed human baseline. Configurable LLM pipelines can thus mimic expert workflows with transparent, inspectable steps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。