提出低成本标签翻转攻击,精准操控大模型对齐方向。
Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
- 将标签翻转攻击建模为带线性约束的凸优化问题,求解最小攻击成本。
- 理论推导出攻击成本上下界,实验证明可降低超50%标签翻转次数。
- 适用于评估大模型对齐阶段的鲁棒性,尤其适合小特征维度场景。
大型语言模型(LLMs)在真实系统中日益广泛应用,其安全性至关重要。尽管已有研究从经验上分析了强化学习人类反馈(RLHF)/直接偏好优化(DPO)对齐过程中的数据投毒攻击,但其理论基础仍不清晰。本文研究在不修改对比输出的前提下,通过翻转偏好标签来引导LLM策略向攻击者目标偏离的最小成本投毒攻击。我们将该问题建模为带有线性约束的凸优化问题,推导出最小攻击成本的上下界。作为理论分析的副产品,我们证明任何现有标签翻转攻击均可通过所提方法后处理,以减少所需标签翻转数量而保持预期投毒效果。实验表明,这种成本最小化后处理能显著降低投毒成本,尤其当奖励模型特征维度相对于数据集规模较小时效果更明显。这些发现揭示了RLHF/DPO流程中的根本脆弱性,并提供了评估其对抗低成本投毒攻击能力的工具。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly deployed in real-world systems, making it critical to understand their vulnerabilities. While data poisoning attacks during RLHF/DPO alignment have been studied empirically, their theoretical foundations remain unclear. We investigate the minimum-cost poisoning attack required to steer an LLM's policy toward an attacker's target by flipping preference labels during RLHF/DPO, without altering the compared outputs. We formulate this as a convex optimization problem with linear constraints, deriving lower and upper bounds on the minimum attack cost. As a byproduct of this theoretical analysis, we show that any existing label-flipping attack can be post-processed via our proposed method to reduce the number of label flips required while preserving the intended poisoning effect. Empirical results demonstrate that this cost-minimization post-processing can significantly reduce poisoning costs over baselines, particularly when the reward model's feature dimension is small relative to the dataset size. These findings highlight fundamental vulnerabilities in RLHF/DPO pipelines and provide tools to evaluate their robustness against low-cost poisoning attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。