arXiv:2603.28281cs.LG2026-03

提出抗数据污染的离线多智能体强化学习方法,提升人类反馈下的鲁棒性。

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback

  • 基于线性马尔可夫博弈建模,设计抗ε比例恶意污染的稳健估计算法
  • 在单一覆盖条件下实现纳什均衡误差为O(√ε),且支持准多项式时间求解
  • 首次系统处理离线多智能体强化学习中的人为数据污染问题,适合安全关键场景

我们研究在强污染模型下,离线多智能体强化学习从人类反馈(MARLHF)中的数据鲁棒性问题:给定包含轨迹-偏好元组的数据集 $D$(每个偏好为表示 $n$ 个智能体偏好的 $n$ 维二值向量),其中最多 $ε$ 比例的样本可能被任意篡改。本文采用线性马尔可夫博弈框架。首先,在均匀覆盖假设下(即感兴趣策略在污染前数据中充分覆盖),提出一个稳健估计器,保证纳什均衡差距为 $O(ε^{1 - o(1)})$。其次,在更困难的单边覆盖情形(仅纳什均衡及其单智能体偏离被覆盖)下,算法实现 $O(ar{ε})$ 的纳什差距。两者虽计算不可行,因此将解概念放松至粗略相关均衡(CCE),在相同覆盖条件下,给出一个准多项式时间算法,其 CCE 差距为 $O(ar{ε})$。据我们所知,这是首个系统处理离线 MARLHF 中对抗性数据污染的工作。

原文摘要 · Abstract (English)

We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset $D$ of trajectory-preference tuples (each preference being an $n$-dimensional binary label vector representing each of the $n$ agents' preferences), an $ε$-fraction of the samples may be arbitrarily corrupted. We model the problem using the framework of linear Markov games. First, under a uniform coverage assumption - where every policy of interest is sufficiently represented in the clean (prior to corruption) data - we introduce a robust estimator that guarantees an $O(ε^{1 - o(1)})$ bound on the Nash equilibrium gap. Next, we move to the more challenging unilateral coverage setting, in which only a Nash equilibrium and its single-player deviations are covered. In this case, our proposed algorithm achieves an $O(\sqrtε)$ bound on the Nash gap. Both of these procedures, however, suffer from intractable computation. To address this, we relax our solution concept to coarse correlated equilibria (CCE). Under the same unilateral coverage regime, we derive a quasi-polynomial-time algorithm whose CCE gap scales as $O(\sqrtε)$. To the best of our knowledge, this is the first systematic treatment of adversarial data corruption in offline MARLHF.

多智能体鲁棒学习离线RL人类反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。