arXiv:2505.21395cs.LG2025-05ICML被引 3

提出新方法,同时保护隐私并抗标签噪声,提升语言模型对齐效果。

Square$χ$PO: Differentially Private and Robust $χ^2$-Preference Optimization in Offline Direct Alignment

  • 用平方损失替代标准对数损失,提升鲁棒性与隐私性。
  • 首次实现单策略集中度下的最优隐私率,且支持中央隐私模式。
  • 理论统一分析,适用于隐私与噪声并存场景,适合安全对齐研究者。

本文从理论上研究在偏好标签污染和隐私保护双重约束下,离线语言模型与人类偏好对齐的问题。为此,提出 Square$χ$PO,仅将 $χ$PO 中的标准对数损失替换为概率上的平方损失。得益于该损失的内在性质,该方法在差分隐私与鲁棒性方面达到当前最优。对于本地隐私模型,Square$χ$PO 是首个在一般函数近似下基于单策略集中度实现最优率的算法;在中央隐私模型下,首次实现对提示(响应)和标签的联合隐私保护。在抗 Huber 标签污染方面,是首个在一般函数近似下具备有意义理论保证的对齐方法。更重要的是,Square$χ$PO 可同时处理隐私与污染问题,且发现隐私与污染的处理顺序存在关键影响。此外,该方法可自然扩展至更一般的偏好模型,并在污染与隐私条件下保持前沿理论性能。所有理论结果基于一个关于带污染与隐私约束的最小二乘回归泛化误差的新界,该结果本身也具有独立研究价值。

原文摘要 · Abstract (English)

In this paper, we theoretically study the offline alignment of language models with human preference feedback, under both preference label corruption and privacy protections. To this end, we propose Square$χ$PO, a simple one-line change to $χ$PO where the standard log-loss is replaced by a new square loss over probability. Thanks to the inherent properties of this new loss, we have advanced the state-of-the-art of differentially private and robust offline direct alignment. Specifically, for the local model of label privacy, Square$χ$PO is the first algorithm that attains an optimal rate based on single-policy concentrability even with general function approximations. It also gives the first result under the central model of privacy protection over both prompts (responses) and labels. On the robustness side against Huber label corruption, Square$χ$PO is the first alignment method that has a meaningful theoretical guarantee under general function approximations. More importantly, Square$χ$PO can address privacy protection and corruption simultaneously, where an interesting separation is observed, implying that the order of privacy and corruption matters. Furthermore, we show that Square$χ$PO can also be easily extended to handle the scenario of the general preference model with state-of-the-art guarantees under corruption and privacy. Last but not least, all of our theoretical guarantees enjoy a unified analysis, building upon a new result on the generalization error bounds of least-square regression under corruption and privacy constraints, which we believe is of independent interest to the community.

隐私对齐鲁棒学习差分隐私语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。