发现大模型验证中信号异质性会限制优化效果,提出按成本分层的智能分配策略。
Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains
- 识别出不同成本区间内预测不确定性质量差异显著,部分区域几乎无法区分错误。
- 提出CST和HGA方法,在复杂场景下提升命中率最高达17个百分点。
- 适合关注大模型资源高效部署与验证策略的研究者和工程团队。
在预算受限的大模型验证系统中,选择性计算机制需决定哪些输出值得进一步验证或人工审核。传统做法是通过统一的不确定性或奖励信号进行在线优化,但本文发现:当信号在不同输入成本分层中存在异质性时,全局优化反而可能失效。研究显示,某些区域的不确定性近乎随机,却集中了大量错误。通过构建局部模型,揭示了全局分配偏差的上限与跨层信号质量差异正相关。为此提出四阶段干预框架(阈值、MP-自适应、MP-分层、成本分层阈值),并设计异质性门控分配(HGA)策略,利用预热可比性测试动态选择全局或分层分配。在MBPP与MATH数据集上,使用Qwen3-8B、LLaMA3-8B和GPT-4o-mini验证,全局在线优化效果不稳定;而CST在高度异质场景下命中率提升最多达17个百分点,HGA则在保留优势的同时避免无效分层。结论建议:在优化共享代理前,应先检验其在不同运行场景下的决策可比性,并以该测试结果决定是否启用结构化专精。
原文摘要 · Abstract (English)
Selective-compute LLM systems decide which outputs merit verification, additional reasoning, tool execution, or human audit under a limited budget. It is natural to expect that stronger online optimization over a shared uncertainty or reward signal should improve these decisions. We take a critical look at this assumption and ask: when does optimizing harder fail because the signal is not decision-comparable across inputs? In budgeted LLM verification, we find that uncertainty quality is heteroskedastic across cost strata: some regions exhibit near-random discriminability while concentrating many errors. Under an explicit local model, we characterize the resulting distortion of global allocation and show that its upper bound scales with cross-stratum signal-quality dispersion. To separate weak signals from optimizer instability and structural mismatch, we introduce a controlled intervention hierarchy: Threshold, MP-Adapt, MP-Strat, and cost-stratified thresholding (CST). We then turn the diagnosis into Heterogeneity-Gated Allocation (HGA), which uses a warm-up comparability test to choose between global and cost-stratified allocation. Across MBPP and MATH using Qwen3-8B, LLaMA3-8B, and GPT-4o-mini, global online adaptation yields inconsistent gains over static thresholding; CST improves hit rate by up to 17 percentage points in strongly heterogeneous settings, while HGA preserves most gains and avoids blind stratification when the partition is not useful. These findings suggest a resource-allocation principle for LLM systems: before optimizing harder over a shared proxy, test whether the proxy is decision-comparable across observable operating regimes, and gate structural specialization on that test.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。