让智能体同时提升任务准确率和工具使用效率,实现双赢。
Towards Pareto-Optimal Tool-Integrated Agents with Pareto Ranking Policy Optimization

- 用动态权重调整优化多目标,自动平衡准确率与效率。
- 在数学推理和多跳问答任务中,表现优于传统方法。
- 适合需要高效实用智能体的工业落地场景。
工具集成语言智能体在解决复杂推理任务方面取得显著进展,但现有对齐方法主要关注最大化任务准确率,忽视了工具使用效率等辅助目标,而这对于实际部署至关重要。为此,我们提出 ParetoPO,一种两阶段多目标优化框架,用于在多个竞争目标下对使用工具的大语言模型进行对齐。第一阶段,ParetoPO 采用超体积引导的动态标量法,根据全局帕累托前沿进展自适应调整奖励权重;第二阶段,以帕累托排序为基础计算优势值,通过支配感知的信用分配机制,促进非占优轨迹的学习。该设计实现了跨多个冲突目标的细粒度、动作级优化。在数学推理和多跳问答任务上的实验表明,相较于静态和启发式基线,ParetoPO 能持续发现具有更优准确率-效率权衡的策略。
原文摘要 · Abstract (English)
Recent advances in tool-integrated language agents have significantly improved their ability to solve complex reasoning tasks. However, existing alignment methods predominantly focus on maximizing task accuracy, while overlooking auxiliary objectives such as tool-use efficiency, which are essential for practical deployment. To address this gap, we introduce ParetoPO, a two-stage multi-objective optimization framework for aligning tool-using large language models (LLMs) under competing objectives. In the first stage, ParetoPO leverages hypervolume-guided dynamic scalarization to adapt reward weights based on global Pareto frontier progress. In the second stage, it replaces scalarized learning signals with Pareto-ranking-based advantage computation, promoting nondominated trajectories through dominance-aware credit assignment. This design enables fine-grained, action-level optimization across multiple conflicting objectives. Experimental results on mathematic reasoning and multi-hop QA tasks show that ParetoPO consistently discovers policies with superior accuracy-efficiency trade-offs compared to static and heuristic baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。