arXiv:2605.02463cs.MAcs.AI2026-05被引 1

提出检测多智能体大模型抗压能力的新方法,发现压力中隐藏可学习结构。

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems

论文配图:When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems
图 1 · 摘自论文原文
  • 构建统计框架CAFE,通过对比预期与实际压力分布识别抗压潜力
  • 五种架构在压力下质量下降1/3,但均显示正向分布差距
  • 适合关注模型韧性与未来可进化性的研究者使用

多智能体大模型常通过分解、辩论、专业化和集成推理解决复杂任务,但现有评估侧重鲁棒性(扰动下的性能保持)。本文提出新问题:语义压力是否暴露可支持未来抗脆弱学习的结构化变化?为此引入CAFE(认知抗脆弱性评估框架),通过建模受控的语义应力期望分布,从多维裁判信号重建架构特异的实际有效应力分布,并在凸应力势下用分布间Jensen间隙进行比较。正间隙不意味着即时性能提升,而是表明观测应力分布发生凸扩张变形,暗示存在可学习的压力结构。在银行风险分析基准上测试五种架构(扁平、层级、辩论、元自适应、集成)发现,所有架构在压力下平均评判质量下降约三分之一,但均呈现正分布Jensen间隙且置信区间高于零。结果表明,即时质量下降可与可检测的抗脆弱兼容应力几何并存。因此,CAFE并非抗脆弱学习器,而是识别何时何地值得应用抗脆弱学习的测量层。

原文摘要 · Abstract (English)

Multi-agent LLM systems are increasingly used to solve complex tasks through decomposition, debate, specialization, and ensemble reasoning. However, these systems are usually evaluated in terms of robustness: whether performance is preserved under perturbation. This paper studies a different question: whether semantic stress exposes structured variation that could support future antifragile learning. We introduce CAFE (Cognitive Antifragility Framework for Evaluation), a statistical framework for detecting antifragility-compatible regimes in multi-agent architectures. CAFE models a controlled expected distribution of semantic stressors, reconstructs an architecture-specific observed effective stress distribution from multi-dimensional judge signals, and compares both distributions using a distributional Jensen Gap under a convex stress potential. A positive gap does not imply immediate performance improvement; instead, it indicates a convex-expansive deformation of the observed stress distribution, suggesting that the architecture exposes learnable stress structure. We evaluate CAFE on a banking-risk analysis benchmark with five multi-agent architectures: flat, hierarchical, debate, meta-adaptive, and ensemble. Across all architectures, semantic stress reduces average judged quality by roughly one third. Yet all architectures exhibit positive distributional Jensen Gaps with bootstrap confidence intervals above zero. These results show that immediate quality degradation can coexist with statistically detectable antifragility-compatible stress geometry. CAFE is therefore not an antifragile learner itself, but a measurement layer for identifying when and where antifragility learning may be worth applying.

多智能体抗脆弱性评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。