用多适配器分歧检测未知区域,提升大模型科学发现的探索能力
Epistemic Uncertainty for Test-Time Discovery
- 构建小规模低秩适配器集合,通过预测分歧衡量认知不确定性
- 在4个基准上3项任务最大奖励提升,解决方案多样性显著提高
- 适合追求真实创新、避免陷入熟悉模式的研究者使用
基于大语言模型的自动化科学发现依赖于识别真正新颖的解。标准强化学习会惩罚高方差突变,导致策略偏向熟悉模式,即使平均奖励上升,最大奖励仍停滞不前。克服此限制需区分未探索区域与内在难题,这要求测量独立适应的权重假设之间的分歧,而非依赖单个网络的置信度。UG-TTT通过在冻结基模型上维护一组小型低秩适配器来解决该问题。每标记的分歧通过集合预测与权重假设间的互信息量化,分离出认知不确定性,识别因覆盖不足导致适配器分歧的位置,而非由问题固有难度引起。该度量作为探索奖励引入策略梯度,引导策略向持续存在适配器分歧的位置移动,这些位置正是真实发现可能发生之处。核范数正则化确保适配器彼此保持差异,从而在整个训练过程中维持探索信号。在四个科学发现基准测试中,UG-TTT在三项任务上提升了最大奖励,显著提高了解的多样性,消融实验也证实正则化对维持该行为至关重要。
原文摘要 · Abstract (English)
Automated scientific discovery using large language models relies on identifying genuinely novel solutions. Standard reinforcement learning penalizes high-variance mutations, which leads the policy to prioritize familiar patterns. As a result, the maximum reward plateaus even as the average reward increases. Overcoming this limitation requires a signal that distinguishes unexplored regions from intrinsically difficult problems. This necessitates measuring disagreement across independently adapted weight hypotheses rather than relying on a single network's confidence. UG-TTT addresses this challenge by maintaining a small ensemble of low-rank adapters over a frozen base model. The per-token disagreement, quantified as the mutual information between ensemble predictions and weight hypotheses, isolates epistemic uncertainty and identifies positions where insufficient coverage leads to adapter divergence rather than intrinsic problem difficulty. This measure is incorporated as an exploration bonus into the policy gradient, directing the policy toward positions where persistent adapter disagreement signals low training coverage, the same frontier where genuine discovery is possible. A nuclear norm regularizer ensures the adapters remain distinct from one another, thereby preserving the exploration signal throughout training. Across four scientific discovery benchmarks, UG-TTT increases the maximum reward on three tasks, maintains substantially higher solution diversity, and an ablation study confirms that the regularizer is essential for sustaining this behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。