arXiv:2606.07678cs.LGcs.AI2026-06中稿 · EMNLP

用几何方法筛选安全对齐数据,少用90%数据仍保持效果

DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment

论文配图:DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment
图 1 · 摘自论文原文
  • 将偏好对视为模型空间中的方向,进行几何分解
  • 仅用11%数据实现接近全量训练的安全性能
  • 无需训练即可快速筛选,适合资源有限的对齐场景

大语言模型的安全对齐依赖偏好数据,但现有流程常使用大量冗余数据。现有数据选择方法通常独立评分每个偏好对,将方向性偏好信息压缩为标量质量或多样性分数,这种样本中心视角在多数据集设置下尤为受限,因共享安全方向与数据集特异性风险并存。我们提出DOG-DPO,一种无需训练的数据选择框架,将偏好对视为结构化的几何信号。首先将每个偏好对表示为模型表示空间中的方向;随后将多数据集偏好几何分解为全局锚定子空间与数据集特异性残差子空间;最后通过最大化基于多样性的覆盖度来选择子集,确保在DPO训练前广泛且非冗余地覆盖对齐方向。在六个安全基准和两个模型主干上,DOG-DPO仅使用11%的偏好对即实现强大的效用-鲁棒性权衡,恢复了全量数据训练的大部分安全收益,且完全无教师、无训练,显著快于代表性选择基线。

原文摘要 · Abstract (English)

Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Existing data selection methods typically score each preference pair independently, collapsing directional preference information into scalar quality or diversity scores. This sample-centric view is especially limiting in multi-dataset settings, where shared safety directions coexist with dataset-specific residual risks. We propose DOG-DPO, a training-free data selection framework that treats preference pairs as structured geometric signals. DOG-DPO first represents each preference pair as a direction in model representation space. It then decomposes multi-dataset preference geometry into a global anchor subspace and dataset-specific residual subspaces. Finally, it selects subsets by maximizing diversity-based coverage, encouraging broad, non-redundant coverage of alignment directions before DPO training. Across six safety benchmarks and two model backbones, DOG-DPO achieves a strong utility-robustness trade-off using only 11% of the preference pairs. It recovers most of the safety gains of full-data training while remaining entirely teacher-free, training-free, and substantially faster than representative selection baselines.

安全对齐数据筛选几何方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。