让AI自动调节生成目标的权重,更好平衡创意与事实。
MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization
- 用元学习动态调整多目标奖励权重,适应不同任务需求。
- 在7个基准上优于固定权重方法,减少无效生成内容。
- 适合需要灵活权衡多个目标的复杂文本生成场景。
组相对策略优化(GRPO)已成为对齐大语言模型的有效范式,但其效果主要局限于具备可验证真实答案的领域。将GRPO拓展至开放域仍面临重大挑战,因为不受约束的生成涉及多维度且常冲突的目标(如创造性与事实性),静态的奖励加权机制本质次优。为此,我们提出MAESTRO(Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization),引入一种元认知协调层,将奖励加权视为动态潜在策略,利用模型终态隐藏状态作为语义瓶颈,感知任务特定优先级。该问题被建模为双层优化框架中的上下文老虎机问题,轻量级导引网络通过组相对优势作为元奖励信号,与策略共同演化。在七个基准测试中,MAESTRO持续优于单奖励和静态多目标基线,同时保持了GRPO的效率优势,在某些场景下甚至减少了冗余生成。
原文摘要 · Abstract (English)
Group-Relative Policy Optimization (GRPO) has emerged as an efficient paradigm for aligning Large Language Models (LLMs), yet its efficacy is primarily confined to domains with verifiable ground truths. Extending GRPO to open-domain settings remains a critical challenge, as unconstrained generation entails multi-faceted and often conflicting objectives - such as creativity versus factuality - where rigid, static reward scalarization is inherently suboptimal. To address this, we propose MAESTRO (Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization), which introduces a meta-cognitive orchestration layer that treats reward scalarization as a dynamic latent policy, leveraging the model's terminal hidden states as a semantic bottleneck to perceive task-specific priorities. We formulate this as a contextual bandit problem within a bi-level optimization framework, where a lightweight Conductor network co-evolves with the policy by utilizing group-relative advantages as a meta-reward signal. Across seven benchmarks, MAESTRO consistently outperforms single-reward and static multi-objective baselines, while preserving the efficiency advantages of GRPO, and in some settings even reducing redundant generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。