揭示多阶段任务中因果归因的极限,指导高效检查点设计
The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning
- 用信息论证明早期步骤信号随深度指数衰减
- 终点数据学习需指数级样本量,深度超过临界阈值即失效
- 提出最优检查点布局,适用于制造与AI推理链监督
制造产线、服务流程、供应链及AI推理链均面临同一挑战:将最终结果归因于导致它的中间环节。我们建立了一个信息论下界,表明早期步骤与最终结果之间的信号随深度呈指数衰减,导致可靠的信用分配需要指数级样本量。我们证明了四个核心结论:第一,信号衰减定理——将结果归因于早期步骤的样本复杂度随中间步骤数呈指数增长;第二,宽度限制:并行推进仅能提供对数级缓解,相关性限制了有效独立样本数量;第三,目标错配:累加奖励优化的是错误目标,而序列正确性要求每一步都准确;第四,最优检查点设计:同质衰减下均匀间隔为极小极大最优,异质衰减下贪心算法可得最优非均匀调度。这些成果共同构建了操作检验与AI监督设计的统一分析基础。
原文摘要 · Abstract (English)
Manufacturing lines, service journeys, supply chains, and AI reasoning chains share a common challenge: attributing a terminal outcome to the intermediate stage that caused it. We establish an information-theoretic barrier to this credit assignment problem: the signal connecting early steps to final outcomes decays exponentially with depth, creating a critical horizon beyond which reliable learning from endpoint data alone requires exponentially many samples. We prove four results. First, a Signal Decay Bound: sample complexity for attributing outcomes to early stages grows exponentially in the number of intervening steps. Second, Width Limits: parallel rollouts provide only logarithmic relief, with correlation capping the effective number of independent samples. Third, an Objective Mismatch: additive reward aggregation optimizes the wrong quantity when sequential validity requires all steps to be correct. Fourth, Optimal Inspection Design: uniform checkpoint spacing is minimax-optimal under homogeneous signal attenuation, while a greedy algorithm yields optimal non-uniform schedules under heterogeneous attenuation. Together, these results provide a common analytical foundation for inspection design in operations and supervision design in AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。