arXiv:2503.13690cs.CLcs.AI2025-03ACL被引 2
用低秩负偏好优化,高效清除大模型敏感内容
Atyaephyra at SemEval-2025 Task 4: Low-Rank Negative Preference Optimization
- 采用低秩适配的负偏好优化方法
- 显著优于基准,实现敏感内容有效移除
- 适合关注模型安全与内容可控性的研究者
我们提交了SemEval-2025任务4的参赛方案,旨在从大语言模型中移除敏感内容。提出一种基于低秩适配的负偏好优化方法,通过该组合高效计算额外正则化项,提升去敏感化过程的稳定性。实验结果表明,该方法显著超越共享任务基准,验证了其有效性。
原文摘要 · Abstract (English)
We present a submission to the SemEval 2025 shared task on unlearning sensitive content from LLMs. Our approach employs negative preference optimization using low-rank adaptation. We show that we can utilize this combination to efficiently compute additional regularization terms, which help with unlearning stabilization. The results of our approach significantly exceed the shared task baselines.
大模型安全负偏好优化低秩适配
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。