用内在统计量实现无需真值的推理扩增,提升复杂任务生成质量。
Intrinsic Selection and Particle Resampling for Inference-Time Scaling Beyond Domain Verifiability

- 基于样本尾部熵设计内在筛选机制,无需外部验证即可评估解质量。
- 在硬数学题上使通过率提升6.1点,工程设计选择准确率提高20%。
- 适用于多模态与专用模型,不依赖奖励模型或精确真值标注。
推理时扩展(ITS)在数学和编程等可验证领域已取得显著成效,依赖低成本验证实现输出选择。然而,将ITS推广至易受系统性错误影响的任务(如初始假设错误或多重约束未满足),通常需昂贵外部求解器或脆弱的模型验证器。本文核心洞见是:并行样本集的内在统计特性,特别是长度调整后的尾部熵,可在无真值的情况下提供鲁棒的解质量判别信号。这些统计量作为难度门控,动态分配计算资源,适配不同扩展模式。首先,内在筛选(iS)后处理排名,跨三个领域表现媲美共识算法,使工程设计选择准确率较pass@1基线提升20%;其次,内在粒子滤波(iPF)扩展至步骤级重采样,引导生成向高置信推理路径演进,在硬数学问题上平均提升pass@1达6.1分;最后,粒子蒸馏(dPF)通过早期对数概率融合与KL引导重采样注入先验指导,帮助绕过系统性推理错误,满足专家评分标准,在复杂临床响应中最高提升26.5%。该流程可无缝应用于通用、领域专用及多模态架构,无需训练奖励模型或精确真值验证,成功将ITS拓展至开放域任务。
原文摘要 · Abstract (English)
Inference-Time Scaling (ITS) has largely succeeded in verifiable domains like math and coding, where cheap verification enables scalable output selection. However, extending ITS to tasks prone to systematic failure - driven by faulty initial assumptions or unmet multidimensional constraints - typically relies on costly external solvers or brittle, model-based verifiers. Our key insight is that the intrinsic statistics of parallel sample sets, specifically length-adjusted tail entropy, provide a robust discriminative signal for solution quality without access to ground truth. Crucially, these statistics serve as a difficulty gate for adaptive compute allocation, dynamically routing problems across scaling regimes. First, Intrinsic Selection (iS) ranks candidates post-hoc, matching consensus-based algorithms across three domains and improving engineering design selection by 20% over pass@1 baselines. Second, Intrinsic Particle Filtering (iPF) generalizes this to step-level resampling, guiding generation toward high-confidence reasoning trajectories to improve pass@1 by 6.1 points on average on hard math problems. Finally, Particle Distillation (dPF) injects privileged guidance via early logit blending and KL-guided resampling, steering generation past systematic reasoning errors to satisfy expert rubrics, yielding up to 26.5% gains on complex clinical responses. Our pipeline applies seamlessly across broad-purpose, domain-specialized, and multimodal architectures, successfully extending ITS to open-ended domains without requiring trained reward models or exact ground-truth verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。