揭示主动学习中性能跃迁的机制根源,解释为何不同策略在不同阶段有效。
A Mechanism-Driven Theory of Phase Transitions in Active Learning

- 将标签预算分段为数据驱动、过渡、模型驱动三阶段,基于主导泛化机制变化定义
- 实验证明策略效率取决于其归纳偏置与当前瓶颈的匹配程度,医学图像也适用
- 首次建立可量化的主动学习动态框架,适合算法设计者和研究者参考
主动学习性能依赖标签预算,但传统以标签数量划分阶段的方法难以跨数据集或模型泛化。本文将预算阶段重新定义为主导泛化机制的转变,通过重构类似PAC的风险成分作为动态交互项,证明机制主导权转移是结构性必然,形成动态泛化瓶颈。通过可测量代理变量与分段回归,提出三阶段分类:数据驱动、过渡、模型驱动。该框架解释了代表性、覆盖度与不确定性策略在不同阶段表现优异的长期现象。在自然图像与医学影像上实验表明,主动学习效率取决于策略归纳偏置与当前瓶颈的对齐程度。此外,自监督表示的迁移更早发生,凸显表示质量对主动学习动态的影响。本工作为下一代感知跃迁的主动学习算法提供了统一理论基础。
原文摘要 · Abstract (English)
Active learning (AL) performance is known to be budget-dependent, yet regimes are typically defined by heuristic label counts that fail to generalize across datasets or architectures. We characterize AL dynamics by reframing budget regimes as shifts in the dominant generalization mechanism. By reinterpreting PAC-style risk components as dynamic interacting terms, we prove that dominance shifts are structurally unavoidable, creating a moving bottleneck for generalization. We operationalize this using measurable proxies and a segmented regression procedure to identify a tripartite taxonomy: data-driven, transition, and model-driven phases. Our framework explains the long-standing observation that representativeness, coverage, and uncertainty strategies excel at different stages. Experiments across natural and medical imaging show that AL efficiency depends on the alignment between the strategy's inductive bias and the active bottleneck. Moreover, self-supervised representation shift transitions earlier along the labeling trajectory, highlighting the role of representation quality in shaping AL dynamics. Overall, this work provides a unified framework for the next generation of transition-aware AL algorithms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。