arXiv:2511.06222cs.CLcs.CY2025-11AAAI被引 3

让大模型先确保安全再追求有用,实现可信与有用的统一

SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization

  • 模型自生成多种回答,自我评估并优化,确保安全优先
  • 在多个基准上提升有用性,同时不降低安全性,优于强基线
  • 适合医疗、法律等高风险场景,可解释且可扩展

在自残、法律或医疗等高风险场景中,大语言模型需兼具可信与有用。但这两者常相冲突。本文提出优先对齐新范式,强制要求‘可信优先于有用’:只有在满足可信阈值(如无害性或诚实性)后,才优化有用性。为此,我们提出自优先对齐(SPA),一种完全无监督框架:模型自生成多样回复,自我评估并迭代优化,通过双准则去噪消除不一致并控制方差。由此构建词典序偏好对,并使用不确定性加权对齐损失进行微调,重点强化高置信度、高差距决策。多基准实验表明,SPA在不牺牲安全性的前提下提升有用性,超越强基线,同时保持通用能力。结果证明SPA是关键大模型应用中可扩展且可解释的对齐策略。

原文摘要 · Abstract (English)

In high-stakes scenarios-such as self-harm, legal, or medical queries-LLMs must be both trustworthy and helpful. However, these goals often conflict. We propose priority alignment, a new alignment paradigm that enforces a strict "trustworthy-before-helpful" ordering: optimization of helpfulness is conditioned on first meeting trustworthy thresholds (e.g., harmlessness or honesty). To realize this, we introduce Self-Priority Alignment (SPA)-a fully unsupervised framework that generates diverse responses, self-evaluates them and refines them by the model itself, and applies dual-criterion denoising to remove inconsistency and control variance. From this, SPA constructs lexicographically ordered preference pairs and fine-tunes the model using an uncertainty-weighted alignment loss that emphasizes high-confidence, high-gap decisions. Experiments across multiple benchmarks show that SPA improves helpfulness without compromising safety, outperforming strong baselines while preserving general capabilities. Our results demonstrate that SPA provides a scalable and interpretable alignment strategy for critical LLM applications.

大模型对齐安全生成自优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。