对比数据混合与模型合并,提升大模型的有用性、诚实性和安全性平衡。
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
- 通过参数级融合解决多目标对齐冲突,优于传统数据混合方法。
- 新方法RESM在3H指标上提升1%-5%,显著增强模型鲁棒性。
- 适合关注AI安全与多目标对齐的研究者和开发者使用。
实现大语言模型在有用性、诚实性和安全性(3H优化)上的平衡是负责任AI的核心。现有数据混合方法依赖专家知识且存在优化信号冲突;而模型合并虽可在参数层面解决冲突,但其在3H优化中的潜力尚未被充分探索。本文首次系统比较了模型合并与数据混合在构建3H对齐大模型中的有效性,揭示了3H维度间的协同与冲突关系,并分析了数据级(data-level)与参数级(parameter-level)方法在缓解冲突方面的优劣。特别地,提出一种新型重加权增强任务单一合并方法RESM,通过异常值加权和稀疏感知秩选择策略,应对3H对齐合并中的偏好噪声累积与层稀疏性适应挑战。大量实验验证了RESM相较于先前数据混合(2%-5%提升)和模型合并(1%-3%提升)方法的有效性与鲁棒性。相关模型已通过Hugging Face公开,供进一步研究。
原文摘要 · Abstract (English)
Achieving balanced alignment of large language models (LLMs) in terms of Helpfulness, Honesty, and Harmlessness (3H optimization) constitutes a cornerstone of responsible AI. Existing methods like data mixture strategies face limitations, including heavy reliance on expert knowledge and conflicting optimization signals. While model merging offers parameter-level conflict-resolution strategies through integrating specialized models' parameters, its potential for 3H optimization remains underexplored. This paper systematically compares the effectiveness of model merging and data mixture methods in constructing 3H-aligned LLMs for the first time, revealing previously overlooked collaborative and conflict relationships among the 3H dimensions and discussing the advantages and drawbacks of data mixture (\textit{data-level}) and model merging (\textit{parameter-level}) methods in mitigating the conflict for balanced 3H optimization. Specially, we propose a novel \textbf{R}eweighting \textbf{E}nhanced task \textbf{S}ingular \textbf{M}erging method, \textbf{RESM}, through outlier weighting and sparsity-aware rank selection strategies to address the challenges of preference noise accumulation and layer sparsity adaptation inherent in 3H-aligned LLM merging. Extensive evaluations can verify the effectiveness and robustness of RESM compared to previous data mixture (2\%-5\% gain) and model merging (1\%-3\% gain) methods in achieving balanced LLM alignment. We release our models through \href{https://huggingface.co/Jinluan}{3H\_Merging} for further investigations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。