arXiv:2605.23956cs.AIcs.LG2026-05被引 2

提出量化复杂AI系统中扰动传播的框架,揭示路径分裂与敏感节点。

QUIVER: A Formal Framework for Quantifying Perturbation Propagation and Bifurcation in Compound AI Systems

论文配图:QUIVER: A Formal Framework for Quantifying Perturbation Propagation and Bifurcation in Compound AI Systems
图 1 · 摘自论文原文
  • 基于图结构定义敏感矩阵与路径分解机制
  • 在8200+追踪中识别出不同架构的敏感特性
  • 适合系统工程师和可靠性研究人员使用

将多个大模型调用串联成有向计算图的复合型AI系统已成为生产环境主流。尽管这些系统包含异构节点与混合输出模式,但现有方法无法量化扰动在其中的传播过程——尤其当节点具有随机性且执行路径可能结构性分叉时。本文提出QUIVER,一种形式化框架,用于测量图结构大模型流水线中的扰动传播。该框架定义:(1) 带类型分派距离度量的敏感矩阵,可分类边为放大器、吸收器或阈值敏感型,并引入出现率提升(occurrence-lift);(2) 轨迹发散分析,将变化分解为数值漂移、结构路径偏离和迭代次数偏差;(3) 分叉阈值,识别引发结构路径改变的最小扰动;(4) 分布保真度,量化节点评估数据集与生产分布的偏离程度。在两个企业级生产流水线和一个公开的DSPy多跳问答流水线上验证,覆盖超过8200个仪器化追踪(32,000+对比较),结果表明:QUIVER能揭示不同架构的差异化敏感特征,区分产生相同发散率但机制不同的级联模式,仅从观测数据预测易发生轨迹分叉的节点,并定位到特定节点字段类别中的陈旧评估数据,而聚合指标无法发现。

原文摘要 · Abstract (English)

Compound AI systems that chain multiple LLM calls into directed computation graphs are now the dominant architecture for production AI. Although these architectures leverage heterogeneous nodes with mixed-mode outputs, no existing framework quantifies how perturbations propagate through such pipelines, where nodes are stochastic and execution paths can diverge structurally. We introduce QUIVER, a formal framework for measuring perturbation propagation in graph-structured LLM pipelines. The framework defines: (1) a sensitivity matrix with type-dispatched distance metrics that classifies edges as amplifiers, absorbers, or threshold-sensitive, complemented by occurrence-lift; (2) trajectory divergence decomposing variation into value drift, structural path divergence, and iteration count divergence; (3) bifurcation thresholds identifying the smallest perturbation that causes structural execution path changes; and (4) distribution faithfulness, quantifying when per node evaluation datasets diverge from production distributions. We validate on two production enterprise pipelines and a public DSPy multihop QA pipeline, three structurally distinct architectures. Across 8,200+ instrumented traces (32,000+ pair comparisons), we demonstrate that QUIVER reveals distinct sensitivity profiles across architectures, distinguishes mechanistically different cascade patterns producing identical divergence rates, predicts nodes prone to trajectory bifurcation from observational data alone, and localizes stale evaluation artifacts to specific node-field categories that aggregate metrics cannot surface.

AI系统扰动传播可靠性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。