提出分阶段推理机制,提升大输出空间下的模型表现
Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces

- 将推理分为广度筛选与精细判断两阶段
- 在多数据集上验证两阶段互补性并提升性能
- 新蒸馏方法优于传统方式,适合复杂标签任务
现代推理模型在需要从数十万到数百万候选标签中选出少数相关标签的挑战性多标签任务上展现出惊人的零样本性能。我们从机制角度探究其原理,将推理过程建模为两个阶段:先广泛筛选候选项,再对结果集进行细粒度推理。我们在多个数据集上提供了证据,证明这两个步骤可分离且具有互补性。基于此机制,我们提出了新的机械式蒸馏策略,该方法在各类任务中均持续优于标准蒸馏。
原文摘要 · Abstract (English)
Modern reasoning models offer surprisingly strong zero-shot performance on challenging multi-label tasks that require selecting a small set of relevant options from hundreds of thousands to millions of candidate labels. We investigate how they achieve this mechanistically. We characterize reasoning as a two-phase process: A broad "shortlisting" of candidates followed by fine-grained reasoning over the resulting set. We provide evidence across a range of datasets that these steps can be isolated and are complementary. Using this characterization, we develop a mechanistic distillation strategy that consistently outperforms standard distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。