arXiv:2606.00329eess.SYcs.LG2026-06被引 1

提出可复现的框架,检验递归系统崩溃预警是否有效。

Benchmarking Recursive-Collapse Warning Claims Under Matched False-Positive Control

论文配图:Benchmarking Recursive-Collapse Warning Claims Under Matched False-Positive Control
图 1 · 摘自论文原文
  • 设计闭环检测框架,监控增益、递归持续性与多样性变化
  • 在两个公开数据集上测试,所有方法均未达可接受预警点
  • 强调在严格误报率约束下,不通过也是科学成果

递归系统可能在显性故障出现前进入类似崩溃的状态——自我强化放大、持续递归和多样性萎缩,掩盖内部退化。我们引入 Loopzero,一个声明受限的基准框架,用于检验递归失败是否遵循特定遥测模式:增益上升(G)、递归持续性(p)和多样性下降(δ)。声明边界在 Lean 中定义;该精益产物不验证真实遥测、基准有效性或检测器性能。我们在两个冻结的公开基准上评估:一个分段公共市场基准(Volmageddon 2018,COVID MWCB 2020)和 MovieLens-25M 离线确定性推荐系统回放。检测器在锁定的等误报率合同下评估(FP ∈ [0.03, 0.07],预先注册),所有配置面临相同告警预算。无论是标准对比方法还是 Loopzero 的预注册分位数检测器,均未达到可接受的操作点。在两个典型基准上,方向性观测一致性成立,但存在邻近时域和行级限制。数字化的 Shumailov 等人(2024)大模型训练轨迹方向上与此模式一致;该领域匹配的误报率评估暂予推迟。贡献在于提供一个可复现、可证伪的基准框架,用于在明确告警预算约束下评估递归崩溃预警声明——不通过本身即为第一类科学成果。

原文摘要 · Abstract (English)

Recursive systems can enter collapse-like regimes -- self-reinforcing amplification, persistent recursion, and narrowing diversity that mask accelerating internal degradation -- before overt failure becomes visible. We introduce Loopzero, a claim-bounded benchmark framework for testing whether recursive failures follow a directional telemetry pattern: rising gain (G), recursive persistence (p), and declining diversity ($δ$). The claim boundary is specified in Lean; the Lean artifact does not verify real telemetry, benchmark validity, or detector performance. We evaluate the bridge on two frozen public-artifact benchmarks: a segmented public-markets benchmark (Volmageddon 2018, COVID MWCB 2020) and a MovieLens-25M offline deterministic recommender replay. Detectors are evaluated under a locked equal-false-positive contract (FP $\in$ [0.03, 0.07], pre-registered) so all configurations face the same alert budget. Neither tested standard comparators nor Loopzero's pre-registered quantile detector achieved an accepted operating point. Directional witness alignment held on both canonical benchmarks, with adjacent-horizon and row-level limitations disclosed. Digitized Shumailov et al. (2024) LLM training-loop trajectories are directionally consistent with the pattern; matched-FP evaluation in that domain is deferred. The contribution is a reproducible, falsifiable benchmark framework for evaluating recursive-collapse warning claims under an explicit alert-budget contract -- non-acceptance reported as a first-class scientific outcome.

递归系统预警框架可复现性误报控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。