arXiv:2511.07086cs.AI2025-11被引 3

用可追溯的LLM流程模拟专家决策,让AI推理过程透明可查。

LLM Driven Processes to Foster Explainable AI

  • 分模块设计,整合敏感性分析、博弈论等三种决策框架。
  • 100次物流测试中,因子对齐率达55.5%,角色匹配率57%。
  • 适合需要可解释性的人工智能应用,如医疗、金融决策支持。

我们提出一种模块化、可解释的LLM代理决策支持系统,将推理过程转化为可审计的中间产物。系统实现三种框架:维斯特的敏感性模型(因子集、带符号影响矩阵、系统角色、反馈回路)、正则形式博弈(策略、收益矩阵、均衡点)以及序贯博弈(角色条件代理、树结构构建、逆向归纳)。每一步均可替换模块。使用默认的GPT-5 LLM配合确定性分析器进行均衡计算与基于矩阵的角色分类,生成可追踪的中间结果而非黑箱输出。在真实物流案例中(100次运行),26个因子的平均因子对齐率为55.5%,运输核心子集达62.9%;角色匹配一致率为57%。采用八项标准评分卡(最高100分)的LLM评判者得分与重构的人类基准相当。因此,可配置的LLM流水线能以透明、可检查的步骤模拟专家工作流。

原文摘要 · Abstract (English)

We present a modular, explainable LLM-agent pipeline for decision support that externalizes reasoning into auditable artifacts. The system instantiates three frameworks: Vester's Sensitivity Model (factor set, signed impact matrix, systemic roles, feedback loops); normal-form games (strategies, payoff matrix, equilibria); and sequential games (role-conditioned agents, tree construction, backward induction), with swappable modules at every step. LLM components (default: GPT-5) are paired with deterministic analyzers for equilibria and matrix-based role classification, yielding traceable intermediates rather than opaque outputs. In a real-world logistics case (100 runs), mean factor alignment with a human baseline was 55.5\% over 26 factors and 62.9\% on the transport-core subset; role agreement over matches was 57\%. An LLM judge using an eight-criterion rubric (max 100) scored runs on par with a reconstructed human baseline. Configurable LLM pipelines can thus mimic expert workflows with transparent, inspectable steps.

可解释AILLM代理决策支持博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。