用马尔可夫动态提前预测多智能体系统错误,减少延迟。
ProMAS: Proactive Error Forecasting for Multi-Agent Systems Using Markov Transition Dynamics
- 通过因果差异特征建模推理过程的语义偏移
- 在Who&When数据集上仅处理27%日志即达22.97%准确率
- 支持实时干预,适合高时效性自主推理场景
将大语言模型引入多智能体系统(MAS)可实现复杂长时任务的协作推理,但其集体智能极易因单一逻辑谬误而迅速传播并导致系统崩溃。现有研究多依赖事后分析,难以实现实时干预。为此,我们提出ProMAS框架,利用马尔可夫转移进行主动误差预测。ProMAS提取因果差异特征以捕捉语义偏移,将其映射到量化向量马尔可夫空间,将推理建模为概率转移过程。通过集成主动预测头与跳跃检测机制,该方法基于风险加速而非静态阈值定位错误。在Who&When基准测试中,ProMAS仅需处理27%的推理日志即可达到22.97%的步级准确率,性能媲美反应式监控器MASC,同时降低73%的数据开销。尽管相比事后方法存在精度妥协,但显著提升干预延迟,平衡了诊断精度与自主推理的实时需求。
原文摘要 · Abstract (English)
The integration of Large Language Models into Multi-Agent Systems (MAS) has enabled the so-lution of complex, long-horizon tasks through collaborative reasoning. However, this collec-tive intelligence is inherently fragile, as a single logical fallacy can rapidly propagate and lead to system-wide failure. Most current research re-lies on post-hoc failure analysis, thereby hinder-ing real-time intervention. To address this, we propose PROMAS, a proactive framework utiliz-ing Markov transitions for predictive error anal-ysis. PROMAS extracts Causal Delta Features to capture semantic displacement, mapping them to a quantized Vector Markov Space to model reasoning as probabilistic transitions. By inte-grating a Proactive Prediction Head with Jump Detection, the method localizes errors via risk acceleration rather than static thresholds. On the Who&When benchmark, PROMAS achieves 22.97% step-level accuracy while processing only 27% of reasoning logs. This performance rivals reactive monitors like MASC while reducing data overhead by 73%. Although this strategy entails an accuracy trade-off compared to post-hoc meth-ods, it significantly improves intervention latency, balancing diagnostic precision with the real-time demands of autonomous reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。