arXiv:2603.13356cs.AIcs.LG2026-03

面对伪装成可信的评估者,新方法通过少量真实审计实现可靠决策。

Learning When to Trust in Contextual Social Bandits

  • 基于审计数据构建每个评估者的上下文信任边界,动态加权反馈。
  • 在80%评估者被攻破时仍能恢复真相,优于均值/中位数等传统鲁棒方法。
  • 理论证明审计频率与性能呈平方根关系,揭示了真实数据的必要性。

稳健强化学习通常假设反馈源要么全局可信,要么在固定预算内被破坏。我们识别出一种更隐蔽的失效模式,称为‘上下文谄媚’:评估者在良性情境下诚实,但在关键情境中系统性偏倚,导致无单一评估者全程可靠,且有害评估者可能在重要情境中形成多数。首个结果是信息论下界:存在两个问题实例,其社会反馈分布完全相同但最优动作不同,证明仅依赖社会反馈的任何算法(包括任意鲁棒聚合器)均产生Ω(T)的隐含后悔。因此,突破上下文谄媚必须依赖额外信息。随后我们证明,以概率p_{aud}可获得的稀疏真实审计足以解决此问题。提出 extsc{ESA},从审计中学习每个评估者的上下文信任边界并相应重加权反馈,其高概率隐含后悔上界为~O(√(T d_{VC}/p_{aud}) + d√T + ε_{tol}T),其中d_{VC}为对手偏倚策略的复杂度。审计依赖项1/√p_{aud}匹配信息论必要性。实验表明,当80%社交层被攻击时, extsc{ESA}仍能恢复真值,而中位数和均值类鲁棒基线失败。

原文摘要 · Abstract (English)

Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global budget. We identify a more subtle failure mode that escapes this dichotomy, which we call \emph{Contextual Sycophancy}. In this failure, evaluators are truthful in benign contexts but systematically biased in critical ones, so that no single evaluator is reliable everywhere and the corrupt evaluators may form a \emph{majority} in the contexts that matter. Our first result is an information-theoretic lower bound. We exhibit two problem instances that induce \emph{identical} social-feedback distributions yet have disjoint optimal actions, proving that \emph{any} algorithm relying on social feedback alone (including any robust aggregator, regardless of breakdown point) incurs $Ω(T)$ latent regret. This shows that breaking contextual sycophancy is impossible without having some information. We then show that a sparse stream of ground-truth audits, available with probability $p_{\mathrm{aud}}$, is sufficient. We propose \ESA, which learns a per-evaluator contextual \emph{trust boundary} from audits and re-weights feedback accordingly, and we prove a high-probability latent-regret bound of $\tilde{\mathcal{O}}\!\big(\sqrt{T\,d_{VC}/p_{\mathrm{aud}}} + d\sqrt{T} + ε_{\mathrm{tol}}T\big)$, where $d_{VC}$ is the complexity of the adversary's bias strategy. The audit-dependence $1/\sqrt{p_{\mathrm{aud}}}$ matches the information-theoretic necessity of audits. Empirically, \ESA\ recovers the ground truth when $80\%$ of the social layer is adversarial, a regime in which median- and mean-based robust baselines fail.

强化学习鲁棒性审计机制社会反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。