多模块大模型组合推理常因局部一致导致全局不一致,本文提出可实时检测与修复的量化方法。
Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agents
- 定义可运行时计算的组合残差eps*,衡量整体推理一致性偏差
- 在1876个组合中33%-94%存在不一致,导致每笔投注额外损失0.115纳特信息
- 提出分层投影修复和在线监控机制,适合高可靠推理系统设计者
多组件大模型代理将概率判断从各自局部视角拼接而成,即使各组件局部一致,组合结果仍可能违反基本概率公理。本文通过组合残差eps*(即输出与联合一致多面体的L2距离)形式化该现象,可基于系统输出与跨组件耦合约束实时计算。产品结构二分法揭示了局部一致是否足以保证全局一致的条件;雷利商预测与实测残差误差小于7%。采用分层博伊尔-迪克斯特拉投影实现确定性修复,引入任意时间有效的e-过程进行序列一致性监控。在包含四个大模型的中等规模面板上,1,876个集成组合中有33%-94%的eps*>0,按比例分配规则计算,1,770次已决赌局平均带来+0.115纳特的后悔损失(若参与者自身能保持一致,则收益降至+0.006)。三种直观缓解策略(检索、分区感知提示、聚合模型)均失效或恶化表现。
原文摘要 · Abstract (English)
Multi-component LLM agents assemble probabilistic claims from components that each see only part of a joint problem; the composition can violate basic probability axioms even when every component is locally coherent. We formalise this locally coherent, globally incoherent failure via the compositional residual eps*, the L2 distance from the composed quote to the joint coherent polytope, computable at runtime from system output and the declared cross-component coupling constraints. A product-structure dichotomy characterises when local coherence suffices, and a Rayleigh-quotient prediction matches the observed residual within 7% on three of four relation classes. A hierarchical Boyle-Dykstra projection repairs the composition deterministically; an anytime-valid e-process gives sequential coherence monitoring. Across 1,876 ensemble cliques on a four-LLM mid-tier panel (frontier-panel rerun in Section 5.5), eps* > 0 on 33-94% of cliques, translating to +0.115 nats per bet of regret on 1,770 resolved bets under the proportional allocation rule (the gain collapses to +0.006 under bettors that themselves coherentise). Three intuitive LLM-side mitigations(retrieval, partition-aware prompting, aggregator-LLM) each fail or regress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。