让大模型更可信:用新方法平衡准确率与自信度。
Less Approximates More: Harmonizing Performance and Confidence Faithfulness via Hybrid Post-Training for High-Stakes Tasks
- 基于渐进推理增益动态调节两种训练方式的权重
- 在少量标注数据下同时提升准确率与判断可信度
- 适合医疗、金融等高风险场景的模型优化
大语言模型在高风险任务中广泛应用,但过度自信的错误推断可能造成严重现实危害,促使“置信度可靠性”问题重回焦点。现有联合优化内部反馈强化学习(RLIF)与推理轨迹引导的推理蒸馏(RD)的方法面临三大挑战:高质量训练语料稀缺、事实性过度自信、以及错误更新被放大。受人类从不确定到确信的认知过程启发,我们提出渐进推理增益(PRG),用于衡量推理步骤是否逐步增强对最终答案的支持。进一步提出HyTuning——一种混合后训练框架,通过类似PRG的指标自适应重分配RD与RLIF的权重,以有限的有监督推理轨迹为稳定锚点,同时利用大量无标签查询实现可扩展性。在多个领域特定与通用基准上的实验表明,该方法在有限监督下既提升了准确率,又实现了置信度可靠性,验证了‘少即更好’的实用效应。
原文摘要 · Abstract (English)
Large language models are increasingly deployed in high-stakes tasks, where confident yet incorrect inferences may cause severe real-world harm, bringing the previously overlooked issue of confidence faithfulness back to the forefront. A promising solution is to jointly optimize unsupervised Reinforcement Learning from Internal Feedback (RLIF) with reasoning-trace-guided Reasoning Distillation (RD), which may face three persistent challenges: scarcity of high-quality training corpora, factually unwarranted overconfidence and indiscriminate fusion that amplifies erroneous updates. Inspired by the human confidence accumulation from uncertainty to certainty, we propose Progressive Reasoning Gain (PRG) to measure whether reasoning steps progressively strengthen support for the final answer. Furthermore, we introduce HyTuning, a hybrid post-training framework that adaptively reweights RD and RLIF via a PRG-style metric, using scarce supervised reasoning traces as a stable anchor while exploiting abundant unlabeled queries for scalability. Experiments on several domain-specific and general benchmarks demonstrate that HyTuning improves accuracy while achieving confidence faithfulness under limited supervision, supporting a practical "Less Approximates More" effect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。