arXiv:2605.12771cs.ROcs.AI2026-05被引 1

动态调节注意力曲线,让机器人在多目标中更稳地找到最优平衡点。

Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization

论文配图:Adaptive Smooth Tchebycheff Attention for Multi-Objective Policy Optimization
图 1 · 摘自论文原文
  • 根据梯度干扰实时调整优化平滑度,实现自适应非凸区域探索
  • 在复杂任务中成功获取线性方法无法达到的非凸帕累托解
  • 适合需要精细权衡多个冲突目标的机器人控制场景

机器人领域中的多目标强化学习需在冲突目标间平衡复杂的非凸权衡。线性加权法虽稳定,但无法获取非凸帕累托前沿上的解;静态非线性加权(如Tchebycheff)理论上可覆盖这些区域,但在深度强化学习中常因梯度方差大而优化不稳。本文提出自适应平滑Tchebycheff框架,通过引入冲突驱动控制器,依据实时梯度干扰动态调节优化景观曲率。当目标趋于一致时,模型可渐进逼近精确的非凸加权;当出现破坏性梯度冲突时,则弹性回退至稳定平滑近似。我们在一个具有挑战性的机器人隐蔽视觉搜寻任务上验证该方法——作为脆弱生态系统监测的代理任务,要求平衡搜索效率、暴露与干扰最小化及探索速度。大量消融实验表明,该冲突感知自适应机制能稳健发现线性基线无法触及、静态非线性方法不稳定的非凸帕累托最优策略。

原文摘要 · Abstract (English)

Multi-objective reinforcement learning in robotic domains requires balancing complex, non-convex trade-offs between conflicting objectives. While linear scalarization methods provide stability, they are theoretically incapable of recovering solutions within non-convex regions of the Pareto front. Conversely, static non-linear scalarizations (e.g., Tchebycheff) can theoretically access these regions but often suffer from severe gradient variance and optimization instability in deep RL. In this work, we propose an Adaptive Smooth Tchebycheff framework that resolves this tension by dynamically modulating the curvature of the optimization landscape. We introduce a novel conflict-driven controller that regulates the optimization smoothness based on real-time gradient interference. This allows the agent to anneal toward precise, non-convex scalarization when objectives align, while elastically reverting to stable, smooth approximations when destructive gradient conflicts emerge. We validate our approach on a challenging robotic stealth visual search task -- a proxy for monitoring of protected/fragile ecosystems -- where an agent must balance search, exposure/interference minimization and exploration speed. Extensive ablations confirm that our conflict-aware adaptation enables the robust discovery of Pareto-optimal policies in non-convex regions inaccessible to linear baselines and unstable for static non-linear methods. Website: https://alejandromllo.github.io/research/pasta/

多目标强化学习机器人控制非凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。