用模型不确定度增强奖励,让大模型决策更高效可靠。
SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards
- 引入多种不确定度指标构建细粒度奖励信号
- 在ALFWorld和WebShop上成功率显著提升
- 适合需要自适应探索的复杂任务系统
大型语言模型(LLMs)越来越多地被用于多步决策任务,有效的奖励设计对引导学习至关重要。尽管已有研究探索了多种奖励塑造和步骤级信用分配方法,但一个关键信号仍被忽视:LLMs的内在不确定性。不确定性反映模型置信度,揭示需探索的位置,并在失败轨迹中提供有价值的学习线索。我们提出SELAUR:一种通过不确定性感知奖励实现自我演化的LLM代理,该强化学习框架将熵、最小置信度和间隔等指标整合为统一的词元级不确定性估计,提供密集的置信度对齐监督;同时采用故障感知奖励重塑机制,将不确定性信号注入步骤级与轨迹级奖励,提升探索效率与学习稳定性。在ALFWorld和WebShop两个基准上的实验表明,该方法在多个强基线上持续提升成功率。消融研究进一步验证了不确定性信号对探索与鲁棒性的增强作用。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed as multi-step decision-making agents, where effective reward design is essential for guiding learning. Although recent work explores various forms of reward shaping and step-level credit assignment, a key signal remains largely overlooked: the intrinsic uncertainty of LLMs. Uncertainty reflects model confidence, reveals where exploration is needed, and offers valuable learning cues even in failed trajectories. We introduce SELAUR: Self Evolving LLM Agent via Uncertainty-aware Rewards, a reinforcement learning framework that incorporates uncertainty directly into the reward design. SELAUR integrates entropy-, least-confidence-, and margin-based metrics into a combined token-level uncertainty estimate, providing dense confidence-aligned supervision, and employs a failure-aware reward reshaping mechanism that injects these uncertainty signals into step- and trajectory-level rewards to improve exploration efficiency and learning stability. Experiments on two benchmarks, ALFWorld and WebShop, show that our method consistently improves success rates over strong baselines. Ablation studies further demonstrate how uncertainty signals enhance exploration and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。